Data processing method and related device

By introducing structured expert knowledge and visual features extracted by visual encoder in the training stage of multimodal large language model, the problem of insufficient image feature extraction capabilities of multimodal large language model is solved, the completeness and noise robustness of the model's input information are improved, and the task execution performance is improved.

CN120236106APending Publication Date: 2025-07-01HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311825055.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing multimodal large language model has limited image feature extraction capabilities, which affects its performance in performing graphic and text inference tasks. It is mainly due to the loss of training knowledge of graphic and text pre-training models, which leads to insufficient extraction capabilities of visual perception modules.

Method used

In the training stage of the multimodal large language model, the structured expert knowledge of the image (first parameter information) is obtained through visual tools and combined with the visual features (second parameter information) extracted by the visual encoder to enrich the model input information and improve the completeness and robustness of the model.

Benefits of technology

By introducing structured expert knowledge, the image information recognition ability and noise robustness of the multimodal large language model are improved, and the performance of the model in tasks such as visual question-and-answer and multimodal dialogue is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236106A_ABST
    Figure CN120236106A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and a related device. The data processing method and the related device can be applied to a training scene of a multi-modal large language model. The method comprises the following steps: acquiring a data set, wherein the data set comprises an image, an instruction and a corresponding label; first parameter information of an image in a data set is obtained through a visual tool, and second parameter information of the image is extracted through a visual encoder. The first parameter information and the second parameter information serve as input samples, the instruction and the label are combined to train the large language model, and the multi-mode large language model is obtained. According to the method, compared with a conventional method that the second parameter information of the image is obtained only through a visual encoder, the completeness of the model input information can be improved through introduction of the first parameter information. Moreover, the first parameter information is added in the training stage of the model, so that the robustness of the model to the noise introduced by the first parameter information can be improved, and the influence of the noise introduced by the first parameter information is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a data processing method and related devices. Background Art

[0002] In recent years, large language model (LLM) technology has developed rapidly. As models and data continue to expand, large language models have shown amazing understanding and reasoning capabilities. However, LLM can only process text-related input data and cannot process visual information. At the same time, the perception ability of the visual base model has also developed rapidly. Through multimodal pre-training of images and texts, visual information and text information can be modally aligned, but the reasoning ability of the visual base model is poor. Therefore, the LLM and the visual perception model are complementarily combined to give birth to the multimodal large language model (MLLM). The multimodal large language model can use LLM to process multimodal information and perform multimodal tasks. The multimodal large language model is more general than the LLM that can only process text information, can process more diverse information, and has good reasoning ability.

[0003] At present, most multimodal large language models use the visual encoder obtained through the image-text pre-training model (such as the Contrastive Language-Image Pre-training (CLIP) model) as the visual perception module, through which the feature information in the image is extracted. Then a simple bridging module is used to combine the visual perception module with the large language model to obtain a multimodal large language model. However, the image-text pre-training model is only trained through a large number of general image-text pairs, and the knowledge contained is lossy, resulting in the limited ability of the multimodal large language model to extract image features, affecting the performance of the multimodal large language model in performing image-text reasoning tasks. Summary of the invention

[0004] The present application provides a data processing method and related devices, which can improve the completeness of input information of a large multimodal language model, thereby improving the performance of the large multimodal language model.

[0005] The first aspect of the present application provides a data processing method that can be applied to a data processing device. The method includes: obtaining a data set, the data set includes an image, an instruction, and a corresponding label; obtaining first parameter information of the image through a visual tool; extracting second parameter information of the image through a visual encoder; using the first parameter information and the second parameter information as input samples, combining the instruction and the label to train a large language model to obtain a multimodal large language model.

[0006] First, obtain the dataset for training the multimodal large language model. The dataset can be an open-source public dataset or a dataset obtained through manual annotation. There is no limitation on the acquisition method of the dataset.

[0007] The dataset can be image-text pairs (image-text pairs), instruction (Instruct) datasets, visual question answering (VQA) datasets, etc. The Instruct dataset includes images, instructions, and the results (labels) obtained based on the images and instructions. The VQA dataset includes images, questions, and the results (labels) obtained based on the images and questions. The image-text pair includes an image and the corresponding text description (label). It can be understood that the questions in the VQA dataset can also be called instructions, and the image-text pair can only be used in the pre-training stage of the multimodal large language model. That is to say, in addition to the image-text pair, the training of the multimodal large language model also requires the Instruct dataset or the VQA dataset. Therefore, the dataset will include images, instructions, and the corresponding labels.

[0008] Then, obtain the first parameter information of the images in the dataset through a visual tool, and obtain the second parameter information of the images through a visual encoder. Among them, the visual tool can also be called an expert model, and the extracted first parameter information is the explicit information of the image, which can be called structured expert knowledge. The visual encoder is a model encoder that needs to be trained and learned, and is used to extract the visual feature information of the image, which is the implicit information of the image.

[0009] Input the first parameter information and the second parameter information of the image, as well as the corresponding instructions (text questions) of the image, into the large language model, and the large language model gives the corresponding results. Then, adjust the parameters of the visual encoder and the large language model according to the loss value between the results and the corresponding true labels, that is, the training of the multimodal large language model is completed. The trained multimodal large language model can be used for tasks such as visual question answering or multimodal dialogue.

[0010] In the first aspect of this application, the first parameter information of the image is obtained through a vision tool, and the second parameter information of the image is obtained through a vision encoder, which improves the completeness of the input information of the model, thereby improving the training effect of the multi-modal large language model and further improving the performance of the multi-modal large language model. Compared with the conventional method of only obtaining the visual feature information (i.e., the second parameter information) of the image through the vision encoder, the introduction of structured expert knowledge (i.e., the first parameter information) can further improve the model's recognition ability for image information. Moreover, although the information obtained through the vision tool (i.e., the first parameter information) will introduce additional noise, adding the first parameter information during the training phase of the model can improve the robustness of the model to the noise introduced by the first parameter information and reduce the impact of the noise introduced by the first parameter information.

[0011] In a possible implementation of the first aspect, the above step: obtaining the first parameter information of the image through a vision tool includes: extracting the objects in the image and the coordinate information of the objects through an object detection tool.

[0012] In this possible solution, the vision tool is an object detection tool. Obtaining the first parameter information through the vision tool specifically means: extracting the objects in the image and the coordinate information of the objects through the object detection tool. Exemplarily, the objects and their corresponding coordinate information can be obtained by combining the recognize anything model (RAM) with Grounding-DINO. The objects and the corresponding coordinate information can be expressed as: (i.e., object 1 and the corresponding coordinates), (i.e., object 2 and the corresponding coordinates)….

[0013] Training the model by extracting the coordinate information of the objects in the image can improve the model's recognition ability for the spatial relationship of the objects, so that the model can have better performance in position-related recognition tasks.

[0014] In a possible implementation of the first aspect, the above step: obtaining the first parameter information of the image through a vision tool includes: extracting the text information in the image through an optical character recognition tool.

[0015] In this possible solution, the vision tool is an optical character recognition tool (OCR). Obtaining the first parameter information through the vision tool specifically means: extracting the text information in the image through the OCR tool. Exemplarily, the text information contained in the image is extracted by EasyOCR. The extracted text information can be expressed as: ocr_str1 (character 1), ocr_str2 (character 2)….

[0016] Training the model by extracting the text information contained in the image can improve the model's ability to recognize OCR text, so that the model can have better performance in text-related recognition tasks.

[0017] In a possible implementation of the first aspect, the above step: extracting the second parameter information of the image through the visual encoder includes: extracting the general semantic information of the image through the general semantic encoder.

[0018] The image semantics is the meaning of the content in the image. The semantics of the image can be divided into high-level semantics, middle-level semantics, and low-level semantics. High-level semantics can also be called general semantics, which is used to represent the overall meaning of the image and is also the thing in the image that is closest to human understanding. Middle-level semantics are different attribute features in the image, and low-level semantics include features such as the contours, edges, colors, textures, and shapes of the objects in the image. Exemplarily, if the image includes a dog holding a ball, the extracted general semantic feature is a dog holding a ball, the middle-level semantic features are the dog and the ball, and the low-level semantic features are the features such as the colors, textures, and shapes of the dog and the ball.

[0019] In this possible solution, the visual encoder is a general semantic encoder. Extracting the second parameter information of the image through the visual encoder specifically means: extracting the general semantic features of the image through the general semantic encoder. Exemplarily, using EVA-CLIP as the general semantic encoder to extract the general semantic knowledge of the image.

[0020] By extracting the general semantic features, the overall content of the image can be obtained. Therefore, using the general semantic features to train the large language model, the trained model can recognize the overall meaning expressed by the image, so that the results given by the model can have higher accuracy.

[0021] In a possible implementation of the first aspect, the above step: extracting the second parameter information of the image through the visual encoder includes: extracting the low-level semantic information of the image through the low-level semantic encoder.

[0022] In this possible solution, the visual encoder is the underlying semantic encoder. The specific process of extracting the second parameter information of the image through the visual encoder is as follows: extracting the underlying semantic features of the image through the underlying semantic encoder. Exemplarily, a vector quantized generative adversarial network (VQ-GAN) is used as the underlying semantic encoder to extract the underlying semantic information of the image. The underlying semantics include features such as the contours, edges, colors, textures, and shapes of the objects in the image. That is to say, more accurate and detailed content of the image can be obtained by extracting the underlying semantic features of the image. Therefore, training a large language model using the underlying semantic features can improve the model's recognition ability for image detail features. Moreover, the input information of the model can be further enriched, thereby improving the training effect of the model.

[0023] In a possible implementation manner of the first aspect, the above step: extracting the second parameter information of the image through the visual encoder includes: extracting the table structure information of the image through the chart structure encoder.

[0024] In this possible solution, the visual encoder is the chart structure encoder. The specific process of extracting the second parameter information of the image through the visual encoder is as follows: extracting the table structure features of the image through the chart structure encoder. Exemplarily, Pix2Struct is used as the chart structure encoder to extract the table structure features of the image. By using the table structure features of the image to train the large language model, the model's recognition ability for document tables can be improved. Moreover, the input information of the model can be further enriched, thereby improving the training effect of the model.

[0025] In a possible implementation manner of the first aspect, the method further includes: using the first parameter information and the second parameter information as input samples, and combining instructions and labels to perform inference on the multimodal large language model.

[0026] The dataset can be divided into a training dataset and a test dataset. After obtaining the first parameter information and the second parameter information of the images in the dataset, the part corresponding to the training dataset is used for model training, and the part corresponding to the test dataset is used for inference on the trained model (i.e., the multimodal large language model). Specifically, the first parameter information and the second parameter information belonging to the test dataset are used as the input of the multimodal large language model, and the multimodal large language model is subjected to inference testing by combining the corresponding instructions and labels in the test dataset.

[0027] In this possible implementation manner, adding the first parameter information in the inference stage can improve the completeness of the data in the inference stage, thereby better improving the model performance.

[0028] The second aspect of this application provides a data processing device, including an acquisition unit, an extraction unit, and a training unit. The acquisition unit is used to acquire a data set, which includes images, instructions, and corresponding labels; the acquisition unit is also used to acquire first parameter information of the images through a vision tool; the extraction unit is used to extract second parameter information of the images through a vision encoder; the training unit is used to use the first parameter information and the second parameter information as input samples, and combine the instructions and labels to train a large language model to obtain a multi-modal large language model.

[0029] In a possible implementation manner of the second aspect, the acquisition unit is specifically used to extract the objects in the images and the coordinate information of the objects through an object detection tool.

[0030] In a possible implementation manner of the second aspect, the acquisition unit is specifically used to acquire the text information in the images through an optical character recognition tool.

[0031] In a possible implementation manner of the second aspect, the extraction unit is specifically used to extract the general semantic information of the images through a general semantic encoder.

[0032] In a possible implementation manner of the second aspect, the extraction unit is specifically used to extract the underlying semantic information of the images through an underlying semantic encoder.

[0033] In a possible implementation manner of the second aspect, the extraction unit is specifically used to extract the table structure information of the images through a chart structure encoder.

[0034] In a possible implementation manner of the second aspect, the device further includes an inference unit, which is used to use the first parameter information and the second parameter information as input samples, and combine the instructions and labels to perform inference on the multi-modal large language model.

[0035] The data processing device provided in the second aspect of this application is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0036] The third aspect of this application provides a data processing device, including a processor and a memory. The memory is used to store instructions, and the processor is used to acquire the instructions stored in the memory to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0037] The fourth aspect of this application provides a computer-readable storage medium, which includes instructions. When the instructions are run on a computer, the computer is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0038] The fifth aspect of this application provides a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the method described in the first aspect or any possible implementation of the first aspect.

[0039] The sixth aspect of this application provides a chip system. The chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instructions to execute the method described in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of a system architecture for the data processing method provided by an embodiment of this application;

[0041] Figure 2 It is a schematic diagram of an embodiment of the data processing method provided by an embodiment of this application;

[0042] Figure 3 It is a schematic diagram of another embodiment of the data processing method provided by an embodiment of this application;

[0043] Figure 4 It is a schematic diagram of the technical effects brought by the data processing method provided by an embodiment of this application;

[0044] Figure 5 It is a schematic diagram of another technical effect brought by the data processing method provided by an embodiment of this application;

[0045] Figure 6 It is a schematic diagram of the structure of a data processing device provided by an embodiment of this application;

[0046] Figure 7 It is a schematic diagram of another structure of the data processing device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The embodiments of this application provide a data processing method, which can improve the completeness of the input information of the multi-modal large language model, thereby improving the performance of the multi-modal large language model. The embodiments of this application also provide corresponding devices, computer-readable storage media, computer program products, etc. The following will be described separately.

[0048] Next, with reference to the accompanying drawings, the embodiments of this application will be described. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Those of ordinary skill in the art can understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0049] The terms "system" and "network" in the specification, claims and the above-mentioned drawings of this application may be used interchangeably. Unless otherwise specified, ordinal numbers such as "first" and "second" are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects. It should be understood that the terms used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0050] For ease of understanding, the relevant terms and concepts mainly involved in the embodiments of this application are introduced below.

[0051] 1. Large Language Model (LLM)

[0052] A large language model, also known as a large-scale language model, is an artificial intelligence model designed to understand and generate human language. LLM is trained on a large amount of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. The characteristic of LLM is its large scale, containing billions of parameters, which helps them learn complex patterns in language data. These models are usually based on deep learning architectures, such as transformers.

[0053] 2. Multi-modal Large Language Model (MLLM)

[0054] A model extended from the large language model with the ability to receive and reason about multi-modal information.

[0055] 3. Contrastive Language-Image Pre-training (CLIP)

[0056] Clip is a pre-trained model that can process text and images simultaneously. The core idea of the Clip model is to improve the model's performance by learning the matching relationship between images and text. Specifically, the Clip model consists of two main components: a convolutional neural network (CNN) for processing images and a Transformer model for processing text. Both components are trained to map the input information into the same embedding space and make the distance between similar images and text closer in the embedding space. The pre-training of the Clip model is divided into two stages: the first stage is to train the Transformer model through a large-scale text dataset so that the model can understand the relationships between texts; the second stage is to use a large-scale image and text dataset to train the entire Clip model so that the model can match the connections between texts and images.

[0057] 4. Vector Quantized Generative Adversarial Network (VQ-GAN)

[0058] VQ-GAN is a generative model based on the Generative Adversarial Network (GAN) that can convert images or text into high-quality images. The VQ-GAN model uses two core parts: Vector Quantized (VQ) and GAN. Among them, VQ is a data compression technology that can represent continuous data as discretized vectors. In VQGAN, the input image or text is mapped to a discretized vector representation in the VQ space. These discretized vectors are then sent to the GAN model for image generation. GAN consists of two models, a generator and a discriminator. The generator is responsible for generating images, and the discriminator is responsible for judging whether the generated images are real images. During the training process, the VQ-GAN model optimizes two loss functions: one for quantization error (i.e., the error between the discretized vector and the continuous value), and the other for the adversarial loss between the generator and the discriminator.

[0059] 5. Pix2Struct

[0060] Pix2Struct is a pre-trained image-text model for pure visual language understanding that can be fine-tuned on tasks containing visual language.

[0061] 6. Recognize Anything Model (RAM)

[0062] RAM is an image tagging model that can identify images of any common category with high precision. RAM provides a new paradigm for the field of image recognition. Using a vast amount of web data without manual annotation, a general model with strong generalization ability can be trained. RAM can automatically identify image tags of more than 6400 categories with relatively high accuracy, spanning academic datasets and commercial products. The development of RAM is divided into four steps. First, image tags are obtained through automatic text semantic parsing. Subsequently, supervised training is carried out by unifying text descriptions and tagging tasks, and an initial model is automatically annotated with the original text and parsed tags as supervision. In the third step, a data engine is used to generate additional annotations and remove incorrect tags. Finally, the model is retrained with the processed data and fine-tuned using a smaller but higher-quality dataset.

[0063] 7.Grounding-DINO

[0064] Grounding DINO is an open-set object detector that combines the Transformer-based detector DINO with ground-truth pre-training. Open-set object detection means that the object categories that can be detected are not limited to the training categories. Grounding DINO can detect any object through human input (such as category names or referential expressions).

[0065] 8.Easy OCR

[0066] Optical character recognition (OCR) refers to a technology for searching, extracting, and recognizing text in pictures. It determines the shape by detecting dark and bright patterns and then translates the shape into computer text using character recognition methods. EasyOCR is an open-source OCR library based on PyTorch that can perform multi-language text recognition. It supports more than 80 languages and has relatively high accuracy and robustness.

[0067] Currently, most multi-modal large language models use visual encoders obtained from image-text pre-training models such as CLIP as visual perception modules to extract feature information from images. Then, a simple bridging module is used to combine the visual perception module with the large language model to obtain a multi-modal large language model. However, image-text pre-training models are only trained with a vast amount of general image-text pairs, and the knowledge contained is lossy, resulting in limited ability of multi-modal large language models to extract image features and affecting the performance of multi-modal large language models in performing tasks.

[0068] In view of this, the embodiments of the present application provide a data processing method, which adds structured expert knowledge obtained from images during the training phase of a multimodal large language model, thereby improving the completeness of the model's input information, and further improving the training effect of the model. Moreover, although the structured expert knowledge will introduce additional noise, adding the structured expert knowledge during the training phase can improve the robustness of the model to noise.

[0069] First, the system architecture of the data processing method provided by the embodiments of the present application will be described in conjunction with Figure 1 the following.

[0070] As Figure 1 shown, the embodiments of the present application obtain the parameter information of the image from two aspects. On the one hand, the visual feature information of the image is extracted through a visual encoder. On the other hand, the structured expert knowledge of the image is further extracted through an expert model. Among them, the visual encoder is the model encoder to be trained, which is used to extract the hidden information of the image. The expert model is a visual tool used to obtain the explicit information of the image. After the feature information of the image is extracted by the visual encoder, it is aligned with the text in modality and spliced together as the visual feature, and then jointly input into the LLM together with the structured expert knowledge and instructions obtained by the expert model. The LLM gives the corresponding result. Among them, the instruction can be a question (Query) or a natural language describing the task. Then, the entire model (including the visual encoder and the LLM) is adjusted according to the difference between the result output by the LLM and the label. The trained model can be used for visual question answering, such as "Why is this picture interesting", or multimodal dialogue and other tasks, such as "Describe this picture in detail".

[0071] Next, please refer to Figure 2 , Figure 2 which is a schematic diagram of an embodiment of the data processing method provided by the embodiments of the present application. As Figure 2 shown, this embodiment includes steps 201 to 205.

[0072] 201. Obtain a data set for training a multimodal large language model.

[0073] Obtain an open-source public dataset, which can be an image-text pair (image-text pair), an instruction (Instruct) dataset, a visual question answering (VQA) dataset, etc. The Instruct dataset includes images, instructions, and results (labels) obtained based on the images and instructions. The VQA dataset includes images, questions, and results (labels) obtained based on the images and questions. The image-text pair includes an image and its corresponding text description (label). It can be understood that the questions in the VQA dataset can also be referred to as instructions, and the image-text pair can only be used in the pre-training stage of the multimodal large language model. That is to say, in addition to the image-text pair, the training of the multimodal large language model also requires the Instruct dataset or the VQA dataset. Therefore, the dataset will include images, instructions, and their corresponding labels.

[0074] Of course, the dataset can also be a dataset obtained through other means instead of an open-source public dataset obtained from the network. For example, the dataset can be obtained through manual annotation. In this embodiment, the acquisition method of the dataset is not limited.

[0075] 202. Obtain the first parameter information of the image through a visual tool.

[0076] Considering that the visual information input into the current multimodal large language model is incomplete, in this embodiment, a multi-expert model is used to obtain richer information (i.e., the first parameter information) of the image. Among them, the expert model can be various visual tools.

[0077] In a possible solution, the visual tool is an object detection tool. Obtaining the first parameter information through the visual tool specifically means: extracting the objects in the image and the coordinate information of the objects through the object detection tool. Specifically, it can be to extract the objects and obtain the coordinate information of the objects by combining RAM and Grounding-DINO, or it can be obtained through object detection (Object Detection). The extraction method of the objects and the coordinate information is not specifically limited here. The objects and their corresponding coordinate information can be expressed as: (i.e., object 1 and its corresponding coordinates), (i.e., object 2 and its corresponding coordinates)….

[0078] Training the model by extracting the coordinate information of the objects in the image can improve the model's recognition ability of the spatial relationship of the objects, so that the model can have better performance in location-related recognition tasks.

[0079] In another possible solution, the visual tool is an optical character recognition tool. Obtaining the first parameter information through the visual tool specifically means: extracting the text information in the image through the OCR tool. For example, extracting the text information contained in the image through EasyOCR. The extracted text information can be expressed as: ocr_str1 (character 1), ocr_str2 (character 2)….

[0080] By training the model by extracting the text information contained in the image, the recognition ability of the model for OCR text can be improved, so that the model can have better performance in text-related recognition tasks.

[0081] In addition, the first parameter information can also be other image-related information, such as the scene graph information generated through the scene graph generation task, etc. The first parameter information is not specifically limited here.

[0082] 203. Extract the second parameter information of the image through the visual encoder.

[0083] In addition to the expert knowledge obtained through the visual tool, in this implementation, the visual feature information (i.e., the second parameter information) of the image is also extracted through the visual encoder of the model to improve the input information of the model as much as possible.

[0084] In a possible solution, the visual encoder is a general semantic encoder. Extracting the second parameter information of the image through the visual encoder specifically means: extracting the general semantic features of the image through the general semantic encoder. Exemplarily, EVA-CLIP distilled using the EVA scheme is used as the general semantic encoder, and the general semantic knowledge of the image is extracted through EVA-CLIP. EVA is a vision-centered foundation model designed to explore the limitations of large-scale visual representations using only publicly accessible data. Based on EVA as the vision model foundation, large-scale trained multi-modal models (such as CLIP) can achieve better performance with fewer samples and less computational effort.

[0085] Image semantics is the meaning of the content in the image. The semantics of the image can be divided into high-level semantics, middle-level semantics, and low-level semantics. High-level semantics can also be called general semantics, which is used to represent the overall meaning of the image and is also the thing in the image that is closest to human understanding. Middle-level semantics are different attribute features in the image, and low-level semantics include features such as the outline, edge, color, texture, and shape of the objects in the image. Exemplarily, if there is a dog holding a ball in the image, the extracted general semantic feature is a dog holding a ball, the middle-level semantic feature is the dog and the ball, and the low-level semantic feature is the color, texture, and shape of the dog and the ball, etc.

[0086] The overall content of an image can be obtained by extracting general semantic features. Therefore, by training a large language model using general semantic features, the trained model can recognize the overall meaning expressed by the image, thereby enabling the results given by the model to have higher accuracy.

[0087] In another possible solution, the visual encoder is a low-level semantic encoder. The specific process of extracting the second parameter information of the image by the visual encoder is as follows: extracting the low-level semantic features of the image through the low-level semantic encoder. For example, using VQ-GAN as the low-level semantic encoder to extract the low-level semantic information of the image. As described above, the low-level semantics include features such as the contours, edges, colors, textures, and shapes of the objects in the image. That is to say, more accurate and detailed content of the image can be obtained by extracting the low-level semantic features of the image.

[0088] In another possible solution, the visual encoder is a chart structure encoder. The specific process of extracting the second parameter information of the image by the visual encoder is as follows: extracting the table structure features of the image through the chart structure encoder. For example, using Pix2Struct as the chart structure encoder to extract the table structure features of the image. By using the table structure features of the image to train the large language model, the recognition ability of the model for document tables can be improved.

[0089] It should be noted that the above several solutions can be combined or used alone. Exemplarily, the visual encoder includes multi-layer perception modules such as a general semantic encoder, a low-level semantic encoder, and a chart structure encoder. First, extract the general semantic features through the general semantic encoder, then extract the low-level semantic features through the low-level semantic encoder, then extract the chart structure features through the chart structure encoder, and finally splice them together as visual features (i.e., the second parameter information) and input them into the large language model. Of course, the visual encoder can also include a general semantic encoder and a low-level semantic encoder, or include a general semantic encoder and a chart structure encoder. And, in addition to the above three solutions, the second parameter information can also be obtained in other ways, such as extracting the middle-level semantic features of the image through a middle-level semantic encoder, etc. In this embodiment, the method for the visual encoder to obtain the second parameter information is not limited.

[0090] It should be understood that the specific implementation of each of the above encoders is only for exemplary illustration, and the functions of the general semantic encoder, the low-level semantic encoder, and the chart structure encoder can also be implemented by other commonly used methods in the industry. For example, using CLIP as the general semantic encoder, which is not specifically limited here.

[0091] 204. Use the first parameter information and the second parameter information as input samples, and combine instructions and labels to train the large language model.

[0092] After obtaining the first parameter information and the second parameter information of the image, the large language model is trained based on the first parameter information and the second parameter information. Here, taking the second parameter information including the general semantic features, underlying semantic features, and chart structure features of the image as an example, the training process will be described.

[0093] Exemplarily, EVA-CLIP is used as the general semantic encoder, VQ-GAN is used as the underlying semantic encoder, Pix2Struct is used as the chart structure encoder, and the large language model uses the open-source LLaMA2-chat(7B). Each visual encoder compresses the visual feature information (i.e., the second parameter information) of the image through a bridging module. Among them, the general semantic encoder uses Querying Transformer (Q-Former) as the bridging module, while the underlying semantic encoder and the chart structure encoder use Perceiver Resampler as the bridging module. The bridging module can compress the visual feature information to a specified size, thereby reducing the computational amount. After compressing the second parameter information through the bridging module, a fully connected layer is used to align it to the text token dimension. A token refers to a basic unit in text, which can usually be a word, a phrase, a punctuation mark, a character, etc., depending on the requirements and methods of text processing.

[0094] In each mini-batch, the images input to the EVA-CLIP branch are scaled to a size of 448*448 pixels, the images input to the VQ-GAN branch are scaled to a size of 256*256 pixels, and the images input to the Pix2Struct branch are scaled to a size of 1024*1024 pixels. Mini-batch is a commonly used training algorithm in machine learning. Mini-batch divides a large dataset into some small datasets, and only one small dataset is used to train the model each time. Generally, the more data in the training dataset, the more accurate the trained model will be. However, if the dataset is too large, it will lead to excessive computational amount and long training time. Therefore, using the mini-batch method can reduce the computational time and memory consumption while ensuring the training accuracy of the model, and bring better generalization performance.

[0095] After obtaining the first parameter information and the second parameter information of the images in the dataset, the first parameter information, the second parameter information, and the instruction text are input into the large language model, and the large language model gives the corresponding results. Then, according to the loss value between the results and the corresponding true labels, the parameters of the vision encoder and the large language model are adjusted, that is, the training of the multi-modal large language model is completed. Specifically, the training of the model can be divided into three stages. The first stage is the pre-training stage. In this stage, the image-text pair data in the obtained dataset is used to train for one epoch, and the base learning rate can be set to 1×10 -4 . One epoch is the process of training the complete training samples once. The second stage is the instruction tuning stage. In this stage, self-instruct training is carried out using the Instruct dataset and the VQA dataset in the dataset and the first parameter information obtained through visual tools, and the base learning rate can be set to 3×10 -5 . The third stage is the specific fine-tuning stage. In this stage, fine-tuning training is carried out on the data of specific tasks, and the base learning rate can be set to 1×10 -5 .

[0096] For multiple vision encoders that obtain the second parameter information, only the general semantic encoder is trained in the first stage of training (i.e., the pre-training stage). And because the pre-training data volume is large and the image-text pairs are not strongly correlated data, the parameters of the large language model are frozen. In the latter two stages of training, the general semantic encoder, the underlying semantic encoder, and the graph structure encoder are all trained, and the large language model is fine-tuned using the method of Low-Rank Adaptation. During the training process, the AdamW optimizer can be used to train the network parameters, and multiple graphics processing units (GPUs) can be used for parallel training to improve the training speed.

[0097] It can be understood that the first stage is the pre-training stage, which requires a large amount of data and a large amount of computation. In this stage, only the second parameter information obtained from the images in the dataset through the general semantic encoder can be used for training. In the second and third stages, both the structured expert knowledge (i.e., the first parameter information) obtained from the images in the dataset through visual tools and the second parameter information obtained from the images in the dataset through multiple vision encoders are used for training.

[0098] In a possible solution, in order to enable the large language model to better understand the first parameter information obtained by the visual tool, the first parameter information is input into the large language model in a standardized manner. Exemplarily, the first parameter information is input in the following format:

[0099] "In addition to the image content, it also provides possible objects contained in the image and their coordinates. / / In addition to the image content, it also provides possible objects contained in the image and their coordinates

[0100] Objects and their coordinates: / /

[0101]

[0102] There may be some OCR text informations in the image. / / There may be some OCR text informations in the image.

[0103] OCR Informations: / / Text information is

[0104] ocr_str1, ocr_str2, ... / / Specific content of text information

[0105] Please combine all the above information when answering the question. / / Please combine all the above information when answering the question.

[0106] It is understandable that, compared with only obtaining image information by the visual encoder, although adding the first parameter information will introduce additional noise, adding the first parameter information during the training phase can improve the robustness of the model to noise.

[0107] After the model is trained, a multimodal large language model is obtained, which can be used for tasks such as visual question answering or multimodal dialogue.

[0108] 205. Use the first parameter information and the second parameter information as input samples, and combine the instructions and the labels to infer the multi-model large language model.

[0109] Optionally, the dataset can be divided into a training dataset and a test dataset. After obtaining the first parameter information and the second parameter information of the images in the dataset, the part corresponding to the training dataset is used for training the model, and the part corresponding to the test dataset is used for inferring the trained model (i.e., the multimodal large language model). Specifically, the first parameter information and the second parameter information to which the test dataset belongs are used as the input of the multimodal large language model, and the multimodal large language model is inferentially tested in combination with the corresponding instructions and labels in the test dataset.

[0110] Adding the first parameter information in both the training stage and the inference stage of the model can improve the training effect and the inference effect of the model, thereby better improving the model performance. Moreover, it has better robustness to the noise introduced by the first parameter information.

[0111] In this embodiment, structured expert knowledge (i.e., the first parameter information) is introduced through a visual tool, which improves the completeness of the model input information and gives full play to the advantages of the expert model. Moreover, the first parameter information is added in the training stage of the model, which can improve the robustness of the model to the noise introduced by the first parameter information and reduce the impact of the noise introduced by the first parameter information. In addition, the input information of the model can be further enriched through a multi-layer perceptron module (i.e., multiple visual encoders), and the recognition ability of the model for image information can be improved. In summary, the data processing method provided in this embodiment can improve the performance of the multimodal large language model.

[0112] The following combines Figure 3 to give a general description of the data processing method provided in the embodiments of the present application.

[0113] As Figure 3 shown, for an image, first, object entities in the image and the coordinate information of the objects are extracted through an object detection tool (such as RAM combined with Grounding-DINO). Exemplarily, if the image includes a dog lying on the floor, the extracted objects are the dog and the floor, where the coordinates of the dog are (0.182, 0.356, 0.702, 0.725), and the coordinates of the floor are (0.0, 0.0, 1.0, 1.0). Moreover, the text information in the image is extracted through an optical character recognition tool (such as EasyOCR). Exemplarily, if the image includes the text "Monday, just Monday", the extracted text information is: Monday, just Monday.

[0114] Then, extract the visual feature information of the image through visual encoding. Specifically, extract the general semantic information of the image through a general semantic encoder (such as EVA-CLIP), extract the low-level semantic information of the image through a low-level semantic encoder (such as VQ-GAN), and extract the diagram structure information of the image through a diagram structure encoder (such as Pix2Struct). The general semantic encoder, the low-level semantic encoder, and the diagram structure encoder all use a bridging module to compress the extracted visual feature information. Among them, the general semantic encoder uses Q-Former as the bridging module, and both the low-level semantic encoder and the diagram structure encoder use Perceiver Resampler as the bridging module. The three visual encoders then align the visual feature information to the text token dimension through fully connected layers respectively.

[0115] Next, take the objects and their coordinate information in the image, as well as the text information included in the image, as structured expert knowledge (i.e., the first parameter information), combine it with the visual feature information output by the visual encoder (i.e., the second parameter information), and the instruction corresponding to the image, such as "Explain why this picture is interesting", and input them into the large language model together. The large language model gives corresponding results according to the input image information and instructions. Then, adjust the parameters of the entire model according to the error between the result and the true label (i.e., the result corresponding to the image and instruction in the dataset), so as to complete the training of the model and obtain a trained multi-modal large language model.

[0116] The following combines Figure 4 and Figure 5 to illustrate the technical effects brought by the data processing method provided in the embodiments of the present application.

[0117] Please refer to Figure 4 , which is the test effect of the multi-modal large language model trained by different methods on the public dataset. As Figure 4 can be seen, the multi-modal large language model trained by the method provided in the embodiments of the present application has better performance on the datasets vqav2, OK-VQA, Text-VQA, VisualMRC, WTQ, DocVQA, and MME.

[0118] Please refer to Figure 5 , which shows the influence of different modules in the embodiments of the present application on the model. As Figure 5As shown in the figure, for the dataset vqav2, the effect obtained by only using a conventional general semantic encoder to extract the visual feature information of the image for training the model is 67.1. The effect is 67.7 when combining the visual feature information extracted by the underlying semantic encoder. The effect is 68.2 when using the general semantic encoder, the underlying semantic encoder, and the chart structure encoder to extract the visual feature information. When training the model with the visual feature information extracted by the above-mentioned multiple visual encoders and combining the first parameter information extracted by the visual tool for inference, the effect is 67.9. When training the model with the visual feature information and the first parameter information extracted by the above-mentioned multiple visual encoders and using the first parameter information for inference, the effect is 70.6. For other dataset effect types, they will not be elaborated here.

[0119] It can be seen from Figure 5 that training the model by extracting visual feature information through more visual encoders can make the model have better performance. When combining the first parameter information in the inference stage, due to the introduction of additional noise by the first parameter information, the performance of the model may not necessarily improve. When combining the first parameter information in the training stage, the robustness of the model to noise can be improved. Therefore, combining the first parameter information in both the training and inference stages can bring a significant improvement to the model performance.

[0120] The above has described the embodiments of the present application from the perspective of the method. Next, the related devices in the embodiments of the present application will be introduced from the perspective of the specific device implementation.

[0121] Please refer to Figure 6 , a schematic diagram of a data processing device 600 is provided in the embodiments of the present application. Among them, the data processing device 600 includes an acquisition unit 601, an extraction unit 602, and a training unit 603.

[0122] The acquisition unit 601 is used to acquire a dataset, and the dataset includes images, instructions, and corresponding labels.

[0123] The acquisition unit 601 is further used to acquire the first parameter information of the image through a visual tool.

[0124] The extraction unit 602 is used to extract the second parameter information of the image through a visual encoder.

[0125] The training unit 603 is used to use the first parameter information and the second parameter information as input samples, and combine the instructions and labels to train a large language model to obtain a multimodal large language model.

[0126] Optionally, the acquisition unit 601 is specifically used to extract the objects in the image and the coordinate information of the objects through an object detection tool.

[0127] Optionally, the obtaining unit 601 is specifically configured to obtain the text information in the image through an optical character recognition tool.

[0128] Optionally, the extraction unit 602 is specifically configured to extract the general semantic information of the image through a general semantic encoder.

[0129] Optionally, the extraction unit 602 is specifically configured to extract the underlying semantic information of the image through an underlying semantic encoder.

[0130] Optionally, the extraction unit 602 is specifically configured to extract the table structure information of the image through a chart structure encoder.

[0131] Optionally, the data processing device 600 further includes an inference unit 604, configured to use the first parameter information and the second parameter information as input samples, and combine instructions and labels to perform inference on the multi-modal large language model.

[0132] Each unit in the data processing device 600 performs the operations of the data processing device in the foregoing Figure 2 and Figure 3 illustrated embodiments, and details are not described herein again.

[0133] Please refer to Figure 7 , which is a possible structural schematic diagram of a data processing device 700 provided by an embodiment of the present application, including a processor 701, a communication interface 702, a memory 703, and a bus 704. The processor 701, the communication interface 702, and the memory 703 are interconnected through the bus 704. In the embodiment of the present application, the processor 701 is used to control and manage the actions of the data processing device. For example, the processor 701 is used to execute Figure 2 the steps performed by the data processing device in the illustrated method embodiment. The communication interface 702 is used to support the data processing device to communicate. The memory 703 is used to store the program code and data of the data processing device.

[0134] Among them, the processor 701 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 704 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is shown in Figure 7 , but it does not mean that there is only one bus or one type of bus.

[0135] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium includes instructions that, when run on a computer, cause the computer to execute the foregoing Figure 2 and Figure 3 methods shown in the embodiments.

[0136] The embodiments of the present application also provide a computer program product containing instructions that, when run on a computer, cause the computer to execute the foregoing Figure 2 and Figure 3 methods shown in the embodiments.

[0137] The embodiments of the present application also provide a chip system. The chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instructions to execute the foregoing Figure 2 and Figure 3 methods shown in the embodiments.

[0138] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0139] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0140] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0141] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0143] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.

Claims

1. A data processing method, characterized in that, Comprising: Obtain a data set, where the data set includes images, instructions, and corresponding labels; Obtain first parameter information of the image through a vision tool; Extract second parameter information of the image through a vision encoder; Use the first parameter information and the second parameter information as input samples, and combine the instructions and the labels to train a large language model to obtain a multi-modal large language model.

2. The method according to claim 1, wherein The obtaining the first parameter information of the image through a vision tool includes: Extract the objects in the image and the coordinate information of the objects through an object detection tool.

3. The method according to claim 1 or 2, characterized in that, The obtaining the first parameter information of the image through a vision tool includes: Extract the text information in the image through an optical character recognition tool.

4. The method according to any one of claims 1 to 3, characterized in that The extracting the second parameter information of the image through a vision encoder includes: Extract the general semantic information of the image through a general semantic encoder.

5. The method according to any one of claims 1 to 4, characterized in that, The extracting the second parameter information of the image through a vision encoder includes: Extract the underlying semantic information of the image through an underlying semantic encoder.

6. The method according to any one of claims 1 to 5, characterized in that The extracting the second parameter information of the image through a vision encoder includes: Extract the tabular structure information of the image through a chart structure encoder.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Use the first parameter information and the second parameter information as input samples, and combine the instructions and the labels to perform inference on the multi-modal large language model.

8. A data processing device, characterized in that, Comprising: An obtaining unit for obtaining a data set, where the data set includes images, instructions, and corresponding labels; The obtaining unit is further configured to obtain first parameter information of the image through a vision tool; An extracting unit for extracting second parameter information of the image through a vision encoder; A training unit for using the first parameter information and the second parameter information as input samples, and combining the instructions and the labels to train a large language model to obtain a multi-modal large language model.

9. The device according to claim 8, characterized in that, The obtaining unit is specifically configured to: Extract the objects in the image and the coordinate information of the objects through an object detection tool.

10. The device according to claim 8 or 9, characterized in that, The obtaining unit is specifically configured to: Obtain the text information in the image through an optical character recognition tool.

11. The device according to any one of claims 8 to 10, characterized in that, The extracting unit is specifically configured to: Extract the general semantic information of the image through a general semantic encoder.

12. The device according to any one of claims 8 to 11, characterized in that The extracting unit is specifically configured to: Extract the underlying semantic information of the image through an underlying semantic encoder.

13. The device according to any one of claims 8 to 12, characterized in that The extracting unit is specifically configured to: Extract the tabular structure information of the image through a chart structure encoder.

14. The device according to any one of claims 8 to 13, characterized in that The apparatus further includes: An inference unit for using the first parameter information and the second parameter information as input samples, and combining the instructions and the labels to perform inference on the multi-modal large language model.

15. A data processing device, characterized in that, Comprising: A processor and a memory; The memory is used to store instructions; The processor is configured to execute the instructions stored in the memory to implement the method according to any one of claims 1 to 7.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by one or more processors, implements the method according to any one of claims 1 to 7.

17. A computer program product comprising instructions, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.