Multi-modal large language model training method, electronic equipment and storage medium
By employing joint training and a multimodal attention mechanism, the fragmentation problem in the training process of large multimodal language models is solved, achieving closer fusion of modal information and stronger multimodal processing capabilities, thereby improving the model's understanding and application performance in complex scenarios.
Patent Information
- Application Number
- CN202510941790.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-28
AI Technical Summary
Existing large-scale multimodal language models suffer from problems such as fragmented training processes, low data utilization efficiency, limited regional localization capabilities, and insufficient integration of multimodal capabilities during training, which affect their processing capabilities and performance in complex multimodal scenarios.
A joint training method is adopted, using some parameters of the pre-trained image-text alignment model as initial parameters and combining them with a comprehensive dataset for joint training. Through multimodal attention mechanism and multi-task learning, information interaction and fusion of data from different modalities are promoted. Semantic segmentation task and progressive training strategy are introduced to improve the model's multimodal processing capability.
It significantly improves the model's overall ability to process multimodal data, enabling it to handle complex tasks containing multiple modal information more naturally and smoothly, and enhancing its performance in various tasks such as text-based question answering, chart understanding, and document summarization generation.
Smart Images

Figure CN120851137A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology applications, specifically to a multimodal large-scale language model training method, electronic device, and storage medium. Background Technology
[0002] Large-scale language models (LLMs) such as ChatGPT, Bloom, and LLaMA have garnered widespread attention for their powerful capabilities in text generation and understanding. These models can be further fine-tuned to align with user intent, demonstrating strong interactive capabilities and the potential to enhance productivity as intelligent assistants. However, LLMs are only applicable to plain text and lack the ability to process image, speech, and video modalities, which significantly limits their application scope. To overcome this limitation, multimodal large-scale language models (MLLMs) such as MiniGPT-4, LLaVA, mPLUG-0w1, and Qwen-VL, which use LLMs as language decoders, aim to enhance the perception and understanding of visual signals. They have demonstrated significant zero-shot capabilities in various open-ended visual and language tasks and exhibited remarkable performance across different domains. These multimodal large-scale language models are trained to align text and images in the first stage of training, and then their generalization ability is improved through instruction tuning in the second stage. Because MLMs lack specific training on visual text understanding datasets, they still face the challenge of understanding the complex relationships between objects in visual text and different types of images, such as charts and documents.
[0003] With the rapid development of artificial intelligence technology, multimodal large-scale language models have demonstrated enormous application potential in various fields such as natural language processing and computer vision. Multimodal large-scale language models aim to integrate data from multiple modalities, including text, images, charts, and documents, to achieve more comprehensive and in-depth information understanding and processing, thereby addressing various complex real-world application scenarios, such as intelligent customer service, intelligent document analysis, and image content understanding and generation.
[0004] In existing multimodal large-scale language model training techniques, a phased training approach is typically adopted. First, the image-text alignment model is trained separately using image-text pair data. The goal of this stage is to enable the model to initially understand the semantic relationships between images and text. By learning from a large amount of image-text pair data, the model can master the correspondence between images and corresponding text descriptions, such as identifying objects in an image and matching them with their corresponding descriptions in the text. However, this phased training approach has significant limitations. Because the image-text alignment model and the subsequent large-scale language model training process are independent, the understanding of image-text relationships accumulated by the image-text alignment model in the early training phase cannot be directly and effectively transferred to the large-scale language model. This results in the model's inability to smoothly integrate information from different modalities, failing to fully leverage the advantages of the image-text alignment model and affecting the overall multimodal processing capabilities of the model.
[0005] After training the image-text alignment model, existing techniques utilize document datasets, chart datasets, table datasets, and natural image datasets to train large-scale language models. While this training method aims to enable the model to understand different types of chart and document data and possess a certain ability to locate image regions, it has revealed many problems in practical applications.
[0006] From a data utilization perspective, various datasets are used separately during training, without fully considering the correlation and complementarity between them. For example, document data may contain rich textual semantic information, while chart data presents data relationships in an intuitive graphical way, tabular data has clear structured information, and natural image data contains rich visual features. However, existing technologies have failed to effectively integrate the advantages of these different types of data, resulting in models being unable to fully utilize the information provided by various data when dealing with complex multimodal scenarios, thus limiting the model's ability to understand complex data.
[0007] Regarding image region localization capabilities, existing technologies claim that models can accurately locate regions within images. However, this localization ability is often based solely on simple image feature extraction. Models may determine region locations by recognizing low-level features such as color, shape, and texture, lacking a deep understanding of the image's semantic information. In practical applications, facing complex semantic scenarios, such as images containing multiple objects with intricate interactions, this simple feature extraction-based localization method often struggles to accurately identify and locate key regions, leading to inaccurate localization results and consequently affecting the model's understanding of the image content.
[0008] Furthermore, existing technologies have shortcomings in multimodal capability fusion. While models can unlock diverse multimodal capabilities, the integration between different modal capabilities is not tight enough. For example, when processing image-text question answering tasks, models may not be able to effectively combine visual information from images with semantic information from text, leading to inaccurate or incomplete answers. When faced with tasks requiring multimodal collaboration, such as document summarization combined with graph understanding, models often perform poorly and fail to fully leverage the advantages of multimodal data.
[0009] Existing training techniques for large multimodal language models suffer from problems such as fragmented training processes, low data utilization efficiency, limited regional localization capabilities, and insufficient fusion of multimodal capabilities. These problems severely restrict the performance and effectiveness of large multimodal language models in practical applications. To address these issues, we propose a training method, electronic device, and storage medium for large multimodal language models. Summary of the Invention
[0010] The purpose of this invention is to provide a multimodal large-scale language model training method, electronic device, and storage medium to solve the problems mentioned in the background art.
[0011] To achieve the above objectives, a method for training a large-scale multimodal language model includes the following steps:
[0012] S1. Construct a comprehensive dataset that includes text-image pairs, document data, chart data, tabular data, and natural image data;
[0013] S2. Use the image-text pairing data to pre-train the image-text alignment model to obtain the pre-trained image-text alignment model;
[0014] S3. Using some parameters of the pre-trained image-text alignment model as initial parameters, and combining them with a comprehensive dataset, a large language model is jointly trained. During the training process, a multimodal attention mechanism is used to promote information interaction and fusion between different modal data, enabling the trained model to deeply understand different types of charts and document data, and to have the ability to accurately locate image regions based on semantics and efficient multimodal collaborative processing capabilities.
[0015] Preferably, during the construction of the comprehensive dataset, when annotating various types of data, the annotation content includes, but is not limited to, the paragraph structure of documents, the title and data meaning of icons, the row and column relationships of tables, and the category and location information of objects in natural images.
[0016] Preferably, when pre-training the image-text alignment model using image-text pair data, a contrastive learning method is adopted to optimize the parameters of the image-text alignment model by maximizing the similarity between image-text pairs and minimizing the similarity between mismatched image-text pairs.
[0017] Preferably, when using some parameters of the pre-trained image-text alignment model as initial parameters and combining them with a comprehensive dataset to jointly train a large language model, a progressive training strategy is adopted. First, a portion of simple comprehensive data is used to initially train the model, and then the complexity and diversity of the data are gradually increased to gradually improve the model's multimodal processing capabilities.
[0018] Preferably, the multimodal attention mechanism includes designing a dedicated query generator for each modality; introducing a modal similarity bias when calculating cross-modal attention weights; and using a hierarchical attention mechanism to process global features and local region features separately.
[0019] Preferably, the method for achieving precise image region localization based on semantics includes: introducing a semantic segmentation task during training, allowing the model to learn to associate regions in an image with corresponding semantic information; and optimizing the loss function so that when locating image regions, the model not only considers the visual features of the regions but also fully considers their semantic meaning.
[0020] As a preferred approach, different loss functions are designed for different types of multimodal tasks during joint training, and these loss functions are optimized simultaneously using a multi-task learning approach.
[0021] Preferably, the different types of multimodal tasks include text-to-image question answering, graph understanding, document summarization generation, and image description generation, with a specific loss function designed for each task.
[0022] An electronic device includes a processor and a memory for storing a computer program, which, when executed by the processor, implements the multimodal large language model training method as described in any of the preceding claims.
[0023] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal large-scale language model training method described above.
[0024] Compared with the prior art, the beneficial effects of the present invention are:
[0025] 1. This invention adopts a method of using some parameters of a pre-trained image-text alignment model as initial parameters and combining them with a comprehensive dataset to jointly train a large language model. This breaks the fragmented training process in the existing technology. Through joint training, the model can more closely integrate different modal information during the training process, make full use of the understanding of image-text relationships in the early training of the image-text alignment model, significantly improve the model's overall ability to process multimodal data, and can handle complex tasks containing multiple modal information more naturally and smoothly.
[0026] 2. This invention constructs a comprehensive dataset containing multiple data types and annotates each type of data, fully considering the correlation and complementarity between different datasets. This comprehensive dataset allows the model to simultaneously access the rich information contained in different modalities during the learning process. For example, the textual semantic information of document data, the intuitive graphical relationships of chart data, the structured information of tabular data, and the visual features of natural image data are integrated, enabling the model to more comprehensively and deeply understand the semantic connotations of multimodal data, thus demonstrating stronger comprehension capabilities when handling complex multimodal scenarios.
[0027] 3. This invention introduces a semantic segmentation task, enabling the model to learn to associate regions in an image with corresponding semantic information. By combining this with an optimized loss function, the model considers not only the visual features of the regions but also their semantic meaning when locating them. This allows for more accurate identification and location of key regions in images, especially in complex semantic scenes with multiple objects and intricate interactions between them. The localization results are more precise, providing strong support for the model's in-depth understanding of image content.
[0028] 4. This invention employs a multimodal attention mechanism and a multi-task learning approach to promote information interaction and fusion among data from different modalities. The multimodal attention mechanism, by designing a dedicated query generator for each modality, introducing modality similarity bias, and using a hierarchical attention mechanism to process global and local features separately, enables the model to dynamically focus on the correlation information between data from different modalities, achieving deep fusion of multimodal information. Simultaneously, specialized loss functions are designed for different types of multimodal tasks and optimized simultaneously using a multi-task learning approach, improving the model's overall performance across various multimodal tasks. This allows the model to perform exceptionally well in tasks such as text-based question answering, chart understanding, document summarization generation, and image description generation, enhancing the model's practicality and generalization ability.
[0029] 5. This invention employs a progressive training strategy, initially training the model using a subset of simple, comprehensive data, and then gradually increasing the complexity and diversity of the data. This training method avoids the problem of the model failing to converge in the early stages of training due to overly complex data, enabling the model to gradually adapt to multimodal data of varying complexity and steadily improve its multimodal processing capabilities. Compared to existing technologies that directly use complex data for training, which may lead to unstable model performance, the progressive training strategy of this invention ensures stable performance improvement during the training process, ultimately resulting in a higher-performing multimodal large-scale language model. Attached Figure Description
[0030] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] See also Figure 1 This invention provides a technical solution: a method for training a multimodal large-scale language model, comprising the following steps:
[0033] S1. Construct a comprehensive dataset that includes text-image pairs, document data, chart data, tabular data, and natural image data.
[0034] It should be noted that the construction of the comprehensive dataset involves collecting text-image pairs, document data, chart data, table data, and natural image data from multiple public data sources and real-world business scenarios. For example, text-image pairs containing charts and corresponding text descriptions are collected from academic paper websites; document data is obtained from various document management systems; professional charts and tables are collected from fields such as finance and scientific research; and natural image data is collected from image databases and actual shooting scenarios. Then, data annotation is performed: for document data, paragraph structure is annotated, such as using specific labels to mark the start and end positions of each paragraph; for chart data, chart titles and the meaning of each data point or data area are annotated; for table data, the row and column relationships are annotated, clarifying the information represented by each row and column; for natural image data, the category and location information of objects in the image are annotated, such as using bounding boxes to mark the location of objects and assigning corresponding category labels to each object. Finally, data integration is performed, combining the annotated data into a comprehensive dataset, followed by data cleaning and preprocessing. The cleaning process includes removing duplicate data and correcting erroneous annotations; the preprocessing process includes normalizing image data and performing word segmentation and word vector conversion on text data.
[0035] S2. Use the image-text pair data to pre-train the image-text alignment model to obtain the pre-trained image-text alignment model.
[0036] It should be noted that in the pre-training of the image-text alignment model, the first step is to select a model, choosing a suitable deep learning model as the basic architecture of the image-text alignment model, such as a Transformer-based model. This model comprises an encoder and decoder structure. The encoder processes image and text inputs, while the decoder generates outputs that match the inputs. Contrastive learning training is then performed. In this training, image-text pairs are first input into the image-text alignment model, and image and text features are extracted separately. Image features can be extracted using models such as Convolutional Neural Networks (CNNs) or Visual Transformers (ViT), while text features can be extracted using word embeddings and a Transformer encoder. The similarity between image-text pairs is then calculated using methods such as cosine similarity. Simultaneously, mismatched image-text pairs are constructed, and their similarity is calculated. Finally, a contrastive learning loss function is defined, optimizing the model's parameters by maximizing the similarity between matching pairs and minimizing the similarity between mismatched pairs. Stochastic gradient descent (SGD) or Adam optimization algorithms are used to update the model parameters. Finally, the model is evaluated and saved. Its performance is assessed on a validation set using metrics such as precision and recall to evaluate its ability to match images and text. When the model performance meets the expected requirements, the pre-trained image-text alignment model parameters are saved.
[0037] S3. Using some parameters of the pre-trained image-text alignment model as initial parameters, and combining them with a comprehensive dataset, a large language model is jointly trained. During the training process, a multimodal attention mechanism is used to promote information interaction and fusion between different modal data, enabling the trained model to deeply understand different types of charts and document data, and to have the ability to accurately locate image regions based on semantics and efficient multimodal collaborative processing capabilities.
[0038] It should be noted that, specifically in the joint training of large-scale language models, the following steps are taken:
[0039] I. Model Initialization: Some parameters of the pre-trained image-text alignment model are used as the initial parameters of the large language model. Specifically, some parameters related to text processing in the image-text alignment model, such as the parameters of the text encoder, can be selected as the initial parameters of the large language model to inherit the advantages of the image-text alignment model in text semantic understanding.
[0040] II. Implementation of Multimodal Attention Mechanism: ① Design a dedicated query generator for each modality (text, image, chart, table). For example, for the text modality, the query generator can generate query vectors based on the semantic information of the text; for the image modality, the query generator can generate query vectors based on the visual features of the image. ② Introduce a modal similarity bias when calculating cross-modal attention weights. By calculating the similarity between features of different modalities, a modal similarity bias term is obtained and added to the attention weight calculation, enabling the model to pay more attention to the information interaction between similar modalities. ③ Use a hierarchical attention mechanism to process global features and local region features separately. For the image modality, a global attention mechanism can be used first to extract the global features of the image, and then a local attention mechanism can be used to focus on key regions in the image; for the text modality, a global attention mechanism can be used first to understand the overall semantics of the text, and then a local attention mechanism can be used to focus on keywords or phrases in the text.
[0041] III. Progressive Training Strategy: ① Initially train the large language model using a small amount of simple, comprehensive data. For example, select text-image pairs and simple document data with small data volume and low complexity for training, allowing the model to initially adapt to the input and processing of multimodal data; ② As training progresses, gradually increase the complexity and diversity of the data. For example, gradually introduce more complex charts, tables, and natural image data, as well as document data containing more semantic information. In the process of increasing data complexity, hyperparameters such as the learning rate can be adjusted appropriately to ensure stable training of the model.
[0042] IV. Multi-task Learning and Loss Function Optimization: ① Define different types of multimodal tasks, such as text-to-image question answering, graph understanding, document summarization generation, and image description generation. For each task, design a specific loss function. For example, for text-to-image question answering, the cross-entropy loss function can be used to measure the difference between the model's predicted answer and the actual answer; for graph understanding, the mean squared error loss function can be used to measure the model's accuracy in understanding graph data. ② During training, simultaneously optimize the loss functions for these different tasks. Multiple loss functions are combined into a single overall loss function through weighted summation. Optimization algorithms are then used to update the model parameters, enabling the model to achieve better performance across different multimodal tasks.
[0043] V. Model Evaluation and Tuning: During training, the model's performance is periodically evaluated on the validation set. Based on the evaluation results, hyperparameters such as learning rate and batch size are adjusted, or the model structure is fine-tuned to further improve the model's multimodal processing capabilities. When the model's performance on the validation set reaches stability and meets the requirements, the trained large-scale language model is saved.
[0044] During the construction of the comprehensive dataset, the annotation of various types of data includes, but is not limited to, the paragraph structure of documents, the title and data meaning of icons, the row and column relationships of tables, and the category and location information of objects in natural images.
[0045] It should be noted that, in data annotation, for document data, the paragraph structure is annotated, such as using specific labels to mark the start and end positions of each paragraph; for chart data, the chart title and the meaning of each data point or data area are annotated; for table data, the row and column relationships of the table are annotated, clarifying the information represented by each row and each column; for natural image data, the category and location information of objects in the image are annotated, such as using bounding boxes to mark the location of objects and assigning corresponding category labels to each object.
[0046] When pre-training the image-text alignment model using image-text pair data, a contrastive learning method is adopted to optimize the parameters of the image-text alignment model by maximizing the similarity between image-text pairs and minimizing the similarity between mismatched images and text.
[0047] It should be noted that the contrastive learning training process involves first inputting image-text pair data into the image-text alignment model to extract image features and text features respectively; then calculating the similarity between image-text pairs using methods such as cosine similarity; simultaneously, constructing mismatched image-text pairs and calculating the similarity between them; finally, defining a contrastive learning loss function to optimize the parameters of the image-text alignment model by maximizing the similarity between matching image-text pairs and minimizing the similarity between mismatched pairs.
[0048] When using some parameters of the pre-trained image-text alignment model as initial parameters and combining them with a comprehensive dataset to jointly train a large language model, a progressive training strategy is adopted. First, a portion of simple comprehensive data is used to initially train the model, and then the complexity and diversity of the data are gradually increased to gradually improve the model's multimodal processing capabilities.
[0049] It's important to note that the progressive training strategy involves the following steps: First, initial training is conducted using a small amount of simple, comprehensive data to train the large language model. For example, training with relatively small datasets and low-complexity text-image pairs and simple document data allows the model to initially adapt to multimodal data input and processing. Then, complexity is gradually increased as training progresses, with the complexity and diversity of the data gradually increased. For example, more complex charts, tables, and natural image data, as well as document data containing more semantic information, are gradually introduced. During the process of increasing data complexity, hyperparameters such as the learning rate can be adjusted appropriately to ensure stable model training.
[0050] The multimodal attention mechanism includes designing a dedicated query generator for each modality; introducing a modality similarity bias when calculating cross-modal attention weights; and using a hierarchical attention mechanism to process global features and local region features separately.
[0051] It should be noted that the multimodal attention mechanism works as follows: First, a dedicated query generator is designed for each modality (text, image, chart, table); then, a modal similarity bias is introduced when calculating cross-modal attention weights. By calculating the similarity between features of different modalities, a modal similarity bias term is obtained and added to the attention weight calculation, enabling the model to pay more attention to the information interaction between similar modalities; finally, a hierarchical attention mechanism is used to process global features and local region features separately.
[0052] An electronic device includes a processor and a memory for storing a computer program, which, when executed by the processor, implements a multimodal large-scale language model training method according to any of the above.
[0053] It should be noted that the processor can be a general-purpose central processing unit (CPU), graphics processing unit (GPU), or a dedicated neural network processor (NPU), used to execute the computer program stored in memory. The memory can be random access memory (RAM), read-only memory (ROM), or solid-state drive (SSD), used to store the computer program, intermediate data generated during training, and model parameters. In use, the electronic device first reads the computer program from memory, which implements the aforementioned multimodal large-scale language model training method. The processor then executes the steps sequentially according to the program's instructions, including constructing a comprehensive dataset, pre-training the image-text alignment model, and jointly training the large-scale language model. During training, the processor continuously reads and writes data from memory, such as input data, model parameters, and loss function values. Once training is complete, the electronic device can save the trained large-scale language model to memory for later use in practical applications.
[0054] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, a multimodal large language model training method that implements any of the above.
[0055] It should be noted that the computer-readable storage medium can be a portable storage device such as an optical disc, USB flash drive, or external hard drive, or a network storage device such as a hard drive on a server. When a user needs to deploy the computer program to an electronic device for model training, the computer-readable storage medium can be connected to the electronic device. The processor of the electronic device will read the computer program on the storage medium and load it into memory for execution. During execution, the processor will complete the training task of the multimodal large language model according to the program's logic, ultimately obtaining the trained model.
[0056] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for training a large-scale multimodal language model, characterized in that, Includes the following steps: S1. Construct a comprehensive dataset that includes text-image pairs, document data, chart data, tabular data, and natural image data; S2. Use the image-text pairing data to pre-train the image-text alignment model to obtain the pre-trained image-text alignment model; S3. Using some parameters of the pre-trained image-text alignment model as initial parameters, and combining them with a comprehensive dataset, a large language model is jointly trained. During the training process, a multimodal attention mechanism is used to promote information interaction and fusion between different modal data, enabling the trained model to deeply understand different types of charts and document data, and to have the ability to accurately locate image regions based on semantics and efficient multimodal collaborative processing capabilities.
2. The method for training a multimodal large-scale language model according to claim 1, characterized in that, During the construction of the comprehensive dataset, the annotation of various types of data includes, but is not limited to, the paragraph structure of documents, the title and data meaning of icons, the row and column relationships of tables, and the category and location information of objects in natural images.
3. The method for training a multimodal large-scale language model according to claim 1, characterized in that, When pre-training the image-text alignment model using image-text pair data, a contrastive learning method is adopted to optimize the parameters of the image-text alignment model by maximizing the similarity between image-text pairs and minimizing the similarity between mismatched images-text pairs.
4. The method for training a multimodal large-scale language model according to claim 1, characterized in that, When using some parameters of the pre-trained image-text alignment model as initial parameters and combining them with a comprehensive dataset to jointly train a large language model, a progressive training strategy is adopted. First, a portion of simple comprehensive data is used to initially train the model, and then the complexity and diversity of the data are gradually increased to gradually improve the model's multimodal processing capabilities.
5. The method for training a multimodal large-scale language model according to claim 1, characterized in that, The multimodal attention mechanism includes designing a dedicated query generator for each modality; introducing a modal similarity bias when calculating cross-modal attention weights; and using a hierarchical attention mechanism to process global features and local region features separately.
6. The method for training a multimodal large-scale language model according to claim 1, characterized in that, The method for achieving precise image region localization based on semantics includes: introducing a semantic segmentation task during training, allowing the model to learn to associate regions in an image with corresponding semantic information; and optimizing the loss function so that when locating image regions, the model not only considers the visual features of the regions but also fully considers their semantic meaning.
7. The method for training a multimodal large-scale language model according to claim 1, characterized in that, During joint training, different loss functions are designed for different types of multimodal tasks, and these loss functions are optimized simultaneously using a multi-task learning approach.
8. The method for training a multimodal large-scale language model according to claim 7, characterized in that, The different types of multimodal tasks include text-to-image question answering, graph understanding, document summarization generation, and image description generation, and a specific loss function is designed for each task.
9. An electronic device, characterized in that, It includes a processor and a memory for storing a computer program, which, when executed by the processor, implements the multimodal large language model training method as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal large language model training method as described in any one of claims 1-8.