Multi-modal large model training method and device
Through the gradual pre-training of multimodal large models and the convergence control strategy of dynamic equalization, the stable training problem caused by the differences between the modals and data imbalance during the training process of the full modal model is solved, and the stable training and optimal performance of the data of each modal are achieved.
Patent Information
- Application Number
- CN202510230447.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
During the training process, due to the differences between the modals and data imbalance between the modals, it is difficult to coordinate, stabilize the training or achieve the best indicators of each modal.
A multimodal large model training method is adopted. The pre-training steps include encoder alignment, graphic knowledge enhancement and full-modal joint training, and the multimodal large model is gradually trained, and the dynamic equalization convergence control strategy and loss weight adjustment are used during the training process to ensure stable training of each mode.
The stable training of data in each mode is realized, and the modal support capabilities of the model are gradually expanded, so that the multimodal large model can achieve better performance under each mode.
Smart Images

Figure CN120068981A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a training method and device for a multi-modal large model. Background Art
[0002] A large model is a general term for large-scale models. Large models usually have a relatively complex network structure and a huge parameter scale, such as hundreds of millions, billions, or tens of billions. Large models are usually models pre-trained with a large number of text samples and have good performance in natural language processing. A multi-modal large model is a model that can process one or more of audio, video, images, and text at the same time and can generate at least one of text, images, and audio at the same time. It is an expansion of the application ability of a large language model on the basis of the large language model. A multi-modal large model that can process various data forms can also be called a full-modal large model.
[0003] However, during the training process of the full-modal model, due to the essential differences and data imbalance between modalities, there may be conflicts and difficulties in coordination when training different modalities and tasks together, making it difficult to train stably or achieve better indicators for each modality. Summary of the Invention
[0004] One or more embodiments of this specification describe a training method and device for a multi-modal large model to solve one or more problems mentioned in the background art.
[0005] According to a first aspect, there is provided a training method for a multi-modal large model, where the multi-modal large model includes: a visual encoder and a visual feature bridging module connected in sequence, an audio encoder and an audio feature bridging module connected in sequence, and a large language model connected to the visual feature bridging module and the audio feature bridging module; the method includes the following pre-training steps: an encoder alignment step: using image-text samples and audio-text samples to train the visual feature bridging module and the audio feature bridging module; a graphic-text knowledge enhancement step: using image-text samples and text samples to train the visual encoder, the visual feature bridging module, and the large language model; a full-modal joint training step: using image-text samples, audio-text samples, and text samples to train the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module, and the large language model.
[0006] In one embodiment, after the pre-training step, the method further includes the following steps of fine-tuning instructions for the multimodal large model: a graphic instruction fine-tuning step: using image text samples and text samples to train the visual encoder, visual feature bridging module and large language model; a visual instruction fine-tuning step: using image text samples and video text samples to train the visual encoder, visual feature bridging module and large language model; an omnimodal instruction fine-tuning step: using image text samples, audio text samples and text samples to train the visual encoder, visual feature bridging module, audio encoder, audio feature bridging module and large language model.
[0007] In one embodiment, the all-modal joint training step includes: in a single parameter update cycle, sampling the same batch of image text samples, audio text samples and text samples in turn; determining the model loss under the corresponding modality by processing a single modality sample in a single batch; and weighting the model losses under various modalities to determine the comprehensive model loss to adjust the corresponding model parameters.
[0008] In one embodiment, in the all-modal joint training step, the model loss determined for the single modality sample is multiplied and balanced with a predetermined loss weight and used to determine the reverse transfer gradient of the model parameters; the loss weight corresponding to the model loss of the single modality is determined in the following manner: when the multimodal large model is trained with a sample set of a predetermined size under the single modality until a first convergence condition is satisfied, a first convergence loss that satisfies the first convergence condition is obtained; the loss weight corresponding to the model loss under the single modality is determined according to the ratio of the inverse of the first model loss to the sum of the inverses of the respective convergence losses under various modalities.
[0009] In one embodiment, the full-modal instruction fine-tuning step includes: during the model parameter adjustment process, executing the following dynamic balanced convergence control strategy: determining the current loss weight under the corresponding mode according to the current slope of the corresponding loss curve under each mode, and the current loss weight under a single mode is negatively correlated with the corresponding current slope; the model loss corresponding to the sample under each mode is balanced by multiplying it with the current loss weight and used for model parameter adjustment.
[0010] In a further embodiment, the current slope of the loss curve under a single modality is determined by: obtaining the verification losses of the previous H verification cycles of the current verification cycle, wherein a single verification cycle includes multiple parameter update cycles under the single modality, and the verification loss of the single verification cycle is determined by the following normalization method: the model loss determined by the verification set for the single verification cycle minus the difference between the minimum value of the loss determined by the verification set for its historical H verification cycles, and the ratio of the difference between the maximum value and the minimum value of the model loss determined by the verification set for the historical H verification cycles;
[0011] Determine the current slope of the loss curve in the corresponding modality in the current validation period according to the slope coefficient of the validation loss at the validation period t fitted by using a linear regression model.
[0012] In another further embodiment, the determining the current loss weight in the corresponding modality according to the current slope of the corresponding loss curve in each modality respectively includes: determining the corresponding convergence scores for the normalized results of the respective current slopes corresponding to each modality, where a single convergence score is negatively correlated with the corresponding current slope; mapping each of the convergence scores to respective importance coefficients through an activation function, and determining the weight coefficients of the model losses in various modalities according to the product of the importance coefficients and the number of modalities; determining the respective current loss weights corresponding to each modality in the current validation period according to the modality weights in various modalities respectively.
[0013] In a further further embodiment, the determining the respective corresponding current loss weights according to the weight coefficients in various modalities respectively includes: for a single modality, determining the current loss weight by using an exponential moving average method of the weight coefficient.
[0014] According to a second aspect, there is provided a training device for a multi-modal large model, where the multi-modal large model includes: a visual encoder and a visual feature bridging module connected in sequence, an audio encoder and an audio feature bridging module connected in sequence, and a large language model connected to the visual feature bridging module and the audio feature bridging module; the device includes:
[0015] An encoder alignment unit configured to train the visual feature bridging module and the audio feature bridging module by using image-text samples and audio-text samples.
[0016] A graphic and text knowledge enhancement unit configured to train the visual encoder, the visual feature bridging module, and the large language model by using image-text samples and text samples.
[0017] A full-modal joint training unit configured to train the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module, and the large language model by using image-text samples, audio-text samples, and text samples.
[0018] In one embodiment, the device also includes an instruction fine-tuning unit, which is configured to perform instruction fine-tuning on the multimodal large model through the following steps: a graphic instruction fine-tuning step: using image text samples and text samples to train the visual encoder, visual feature bridging module and large language model; a visual instruction fine-tuning step: using image text samples and video text samples to train the visual encoder, visual feature bridging module and large language model; an omnimodal instruction fine-tuning step: using image text samples, audio text samples and text samples to train the visual encoder, visual feature bridging module, audio encoder, audio feature bridging module and large language model.
[0019] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.
[0020] According to a fourth aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0021] Through the apparatus and method provided in the embodiments of this specification, a training scheme for a large multimodal model is provided, and each network module in the large multimodal model can be decoupled according to function, and the decoupled network modules can be progressively trained in stages, gradually expanding the model's modal support capabilities and achieving better performance in each modality. This training method can effectively address the problem of large differences in data distribution between modalities and achieve stable training of data in each modality.
[0022] In a further embodiment, the technical problem of unbalanced data volume of each modal training sample can be solved by the technical concept of balanced sampling according to the number of iteration steps, and the technical problem of inconsistent convergence speed of each modal training sample can be solved by calculating the slope of the loss curve of each modality to measure the convergence speed of the modality and dynamically adjusting the loss weight of each modality according to the convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0024] Figure 1 A schematic diagram of a specific implementation architecture of a multimodal large model is shown;
[0025] Figure 2 A schematic diagram of the training process of a multimodal large model according to an embodiment of the present specification is shown;
[0026] Figure 3 shows a schematic diagram of the pre-training process in the training process of a multi-modal large model corresponding to the embodiments of this specification;
[0027] Figure 4 shows a schematic diagram of the instruction fine-tuning process of a multi-modal large model according to an embodiment of this specification;
[0028] Figure 5 shows a schematic diagram of the process for determining the model loss and weights during the instruction fine-tuning process in the embodiments of the specification;
[0029] Figure 6 shows a structural block diagram of a training device for a multi-modal large model according to an embodiment of this specification. Detailed implementation manners
[0030] The following describes the solution provided in this specification with reference to the accompanying drawings.
[0031] Figure 1 shows a schematic diagram of a specific implementation architecture of a multi-modal large model. As Figure 1 shown, the large model LLM is a pre-trained large language model. In the case of using the large language model to process multi-modal data, it can be extended to a multi-modal large model. During the extension process, based on various modal data, front-end and post-processing architectures can be added to the original large model. The data modalities can include, for example: image-text modality, text modality, audio-text modality, video-text modality, etc. Among them, the image-text modality can include an image and one of the following: image description information, text information in the image, question-and-answer information for the image, etc. The audio-text modality can include audio and text data, and the video-text modality can include video and text data.
[0032] The specific multi-modal large model architecture can include: Vision Encoder, Audio Encoder, MLP (Perceptron) linear alignment network, large language model (LLM), Image Generator, Audio Decoder. It can be understood that the large language model itself is pre-trained based on text. Therefore, Figure 1Networks such as BPE encoding, word embedding, semantic management (Language Head), and BPE decoding in [model name] can be modules corresponding to the large language model itself for processing text. The visual encoder and audio encoder respectively correspond to a visual feature bridging module (MLP) and an audio feature bridging module (MLP) for aligning visual features and audio features with the feature space of the large language model. Among them, the visual encoder and the visual feature bridging module connected to the visual encoder can process text-image modality data and video modality data, and the audio encoder and the audio feature bridging module connected to the audio encoder can process audio modality data.
[0033] As Figure 1 shown, the input data of the multi-modal large model can include various modality data such as text, text-image, video, and audio. Each modality data is encoded by the corresponding model encoder and then aligned to a unified feature space via the corresponding feature alignment network (such as Figure 1 each MLP in [model name]) and concatenated into an encoded tensor. Here, the unified feature space can include the space defined by at least one of the vector dimension and the value range. After the encoded tensor is input into the large language model for unified processing, a decoded tensor is obtained. Through the decoders corresponding to various modalities, at least one of text, image, and audio can be decoded.
[0034] Referring to Figure 1 the data processing process of the multi-modal large model shown, it is easy to understand that the technical problems that may be encountered during the training of the multi-modal large model include: stable training problems under various modality data; uneven numbers of training samples for various modalities; inconsistent convergence speeds for various modalities, and so on.
[0035] To enable the large model to be stably trained, this specification provides a technical concept of progressively training the multi-modal large model through various modality samples, so that the multi-modal large model can be stably trained under various modalities. Specifically, for the training samples in various modality mixing situations, some modality processing structures can be frozen first, and only the parameters in the other part of the modality processing structures are adjusted, and then the various modality processing structures are mixed and trained.
[0036] It can be understood that the large language model itself has been pre-trained with a large number of natural language samples and has strong text information processing capabilities. The information content in audio data is usually converted into text information for processing. Therefore, during the training process of the multi-modal large model, bridge modules (visual feature bridge module, audio feature bridge module) that align audio data and image data with the feature data received by the large language model can be trained first. Since the information content of audio data can be converted into text form, it is relatively easy to adapt to text. After training the bridge modules for the graphic and audio modalities so that visual features and audio features can be aligned with the feature space of the large language model, the parameters of the visual encoder, visual feature bridge module, and large language model can be further adjusted using graphic-modal sample data and text-modal sample data to enhance the graphic processing ability of the multi-modal large model architecture. Furthermore, various modal sample data can be used to adjust the parameters of each module of the entire multi-modal large model, enabling the large model to have the ability to process various modal data.
[0037] Through this progressive training concept, the multi-modal large model can be stably trained. Additionally, considering the training efficiency and the accuracy of the multi-modal large model, in some embodiments, the training process can be divided into two stages. First, pre-training is performed using a large number of samples according to the method described above, and then instruction fine-tuning training is performed using fewer samples. Instruction fine-tuning (also known as supervised fine-tuning) is one of the important methods to enhance or activate specific capabilities of the large language model (such as instruction-following ability) after pre-training. In the instruction fine-tuning stage, first focus on the network architecture of the processing part of the graphic-modal data to enhance the graphic processing ability of the multi-modal large model, and then perform joint training on all modal data to improve the model processing accuracy.
[0038] The technical concept of this specification will be described in detail below with reference to the accompanying drawings.
[0039] Figure 2 The training process of a multi-modal large model according to an embodiment of this specification is shown. The execution subject of this process can be any computer, device, or server with certain computing capabilities. Among them, the multi-modal large model here is based on a large language model and also includes a visual encoder, a visual feature bridge module, an audio encoder, and an audio feature bridge module, which are respectively used to process graphic-modal data and audio-modal data.
[0040] Among them, the visual encoder can be various network structures implemented using convolutional neural networks, etc. For example, Figure 1 The visual encoder shown, for example, can be implemented by NaViT and can process pictures of any size. The audio encoder can be implemented by SAN-M, etc. Generally, in the text generation stage, the multi-modal large model can only consider the text part of the audio information.
[0041] The visual encoder and the audio encoder can be pre-trained modules. The visual feature bridging module and the audio feature bridging module can convert the encoded representations determined by the corresponding encoders into features that can be processed by the large language model through linear transformation, such as making the value of a single element fall within a predetermined range (such as greater than 0 and less than 50, etc.), transforming the dimension into a predetermined dimension (such as 100 dimensions), and so on.
[0042] The training samples of the multi-modal large model can include the following modalities: Text, Image, Vision, Audio. Among them, the training samples of various modalities can also correspond to the label data of the corresponding modalities. It can be understood that in the case of using the multi-modal large model to process the image-text sample (which can include images and related text data) to generate pictures, the picture description information text can be generated first, and then the picture can be generated using an ImageGenerator as shown in Figure 1 . Therefore, the label data of the image-text sample can be the picture description information of the picture to be generated in text form. Similarly, in the case of using the multi-modal large model to process the audio-text sample (which can include audio data and can also include text data) to generate audio information, the text information in the audio can also be generated first, and then the audio can be synthesized by adding sound features through an audio decoder. Therefore, the label data of the audio-text sample can also be the text information in the audio. That is to say, the label data of various modal samples can all correspond to text information. In this way, in the process of processing the multi-modal large model, the large language model can perform text generation in an autoregressive manner.
[0043] As Figure 2 shown, the training process of the multi-modal large model can include the following pre-training steps: Step 201, encoder alignment step: Using the image-text sample and the audio-text sample, train the visual feature bridging module and the audio feature bridging module; Step 202, image-text knowledge enhancement step: Using the image-text sample and the text sample, train the visual encoder, the visual feature bridging module and the large language model; Step 203, full-modal joint training step: Using the image-text sample, the audio-text sample and the text sample, adjust the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module and the large language model.
[0044] Each step in the training process of the multi-modal large model corresponds to Figure 3 Phase 1, Phase 2, Phase 3 in Figure 3 . Among them, the part with the flame pattern in Figure 3Describe each step in the training process of the multimodal large model.
[0045] First, in step 201, use image-text samples and audio-text samples to train the visual feature bridging module and the audio feature bridging module.
[0046] It can be understood that the image-text samples correspond to the training samples in the image-text modality, and the input data can at least include images. In some samples, the input data can also include at least one of the text information recognizable in the image, image description information, question-and-answer information about the image (such as asking "Is this a puppy?" and answering "Yes", etc.). Generally, the image-text samples can correspond to text information as the image-text label, which is used to compare with the text generated by the large language model. Similarly, the audio-text samples can also correspond to the initial audio as the input data and the text information as the label data.
[0047] Reference Figure 3 As shown in stage 1 in , at this time, the model parameters in the Vision Encoder, Audio Encoder, and large language model LLM can be frozen for the training of the multimodal large model. The so-called freezing means locking the current value. At this time, it can be considered that the undetermined parameters in the multimodal large model only include the model parameters in the visual feature bridging module and the audio feature bridging module.
[0048] For the audio-text samples, the initial audio can be used as the input data of the multimodal large model. The audio features are extracted by the audio encoder of the multimodal large model, the audio features are adapted and aligned with the input of the large language model by the audio feature bridging module, and after being processed by the large language model, they are decoded to generate the audio text. The generated audio text is subjected to autoregressive training with the label data of the audio-text samples, and the loss corresponding to the audio-text samples can be calculated and determined, such as denoted as the audio loss.
[0049] The input data of the image-text samples are used to extract visual features by the visual encoder of the multimodal large model, the visual features are adapted and aligned with the input of the large language model by the visual feature bridging module, and after being processed by the large language model, they are decoded to generate the text information describing the image to be generated. The text information describing the image to be generated can generate the result picture when processed by the image generator. The image generator can be implemented by a conventional generator. Here, in the process of determining the model loss, the text information output by the multimodal model can be subjected to autoregressive training with the label data of the image-text samples, and the obtained loss is, for example, denoted as the image-text loss.
[0050] It should be noted that the input data of image text samples and audio text samples may both involve text parts, such as image description information, text recognized in images, natural language converted from audio data, and so on. This part can be embedded through the text processing module of the large language model (such as Figure 1 the word embedding in
[0051] and then processed by the large language model together with the alignment features of the image encoding representation and the alignment features of the audio encoding representation.
[0052] In the current training stage, the model loss in a single parameter update cycle can include at least one of the audio loss and the image-text loss. In a single parameter update cycle, the model loss can be determined using the training samples of the current batch (batch), and the gradients of the undetermined parameters in the corresponding bridging module can be determined using the model loss. Thus, using a gradient update method such as the gradient descent method, the undetermined parameters in the corresponding alignment network can be adjusted towards the model loss.
[0053] In one embodiment, if the visual feature bridging module and the audio feature bridging module are trained separately, then in a single parameter update cycle, the samples of a single batch can be image text samples or audio text samples. In this way, the gradients of the undetermined parameters in the visual feature bridging module or the audio feature bridging module can be determined using the model loss.
[0054] In another embodiment, if the visual feature bridging module and the audio feature bridging module are trained together, then in a single parameter update cycle, the image text samples and the audio text samples can be processed in different batches, and the model losses of each batch can be accumulated. In this way, the gradients of the undetermined parameters in the visual feature bridging module and the audio feature bridging module can be determined using the model loss.
[0055] The training of the visual feature bridging module and the audio feature bridging module can be stopped when a predetermined condition is met. The predetermined condition can be, for example: the number of image text samples and audio text samples used reaches a predetermined number, such as 100,000 each; the parameter update cycle reaches a predetermined number of cycles, such as 1000; the model losses in multiple consecutive parameter update cycles are all less than a predetermined value, such as 2; the gradients in multiple consecutive cycles are all less than a predetermined value, such as 0.1; and so on.
[0055] Through the training with a large number of image text samples and audio text samples, the features output by the visual feature bridging module and the audio feature bridging module can be quickly aligned with the large language model, thus laying a solid foundation for the subsequent training of the multi-modal large model.
[0056] Then, through step 202, the visual encoder, the visual feature bridging module, and the large language model image text samples are trained using the image text samples and the text samples.
[0057] Through the training of the visual feature bridging module and the audio alignment network, the encoded features of the image-text modality and the audio modality can be basically aligned with the input feature space of the large language model. Considering the complexity of image data and its large difference from text information, at this time, referring to Figure 3 Stage 2 in, the large model can also be unfrozen, and the model parameters in the visual encoder, the visual feature bridging module, and the large language model can be adjusted to enhance the image-text processing ability of the large model.
[0058] Specifically, image-text samples and text samples can be processed by the multi-modal large model, and the model loss can be determined according to the processing results. The model parameters in the visual encoder, the visual feature bridging module, and the large language model can be adjusted in the direction of reducing the model loss. Among them, during the processing of image-text samples and text samples, a single batch can process data of one modality. In a single parameter update cycle, only one modality sample data can be used or two modality data can be used in multiple batches to determine the current model loss.
[0059] In this way, after the preliminary training of the visual feature bridging module, by introducing text samples and adjusting the networks related to the image-text modality together with the image-text samples, and fine-tuning the model parameters of the large language model, on the basis of ensuring the text processing ability of the large language model, the processing ability of the large language model for image-text modality data can be further improved, laying a foundation for the training of the processing ability of mixed modality data.
[0060] Next, in step 203, the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module, and the large language model image-text sample audio text sample are trained using image-text samples, audio-text samples, and text samples.
[0061] After the large language model has a certain image-text comprehensive processing ability, it is also necessary to further conduct comprehensive training on various modality data so that the multi-modal large model has the corresponding processing ability for any modality data input at any time. Therefore, in this step 203, the multi-modal large model can be trained for full-modal adaptation using image-text samples, audio-text samples, and text samples.
[0062] Referring to Figure 3 the schematic of Stage 3 in, in this stage, all the networks in the multi-modal large model and the model parameters of the large language model are unfrozen, that is, they are all updated according to the model loss.
[0063] Considering that in the actual training process, text samples are relatively easy to collect, while the number of image-text samples and audio-text samples is small, especially for audio-text samples, which are difficult to collect, resulting in a large difference in the number of samples between different modalities. There are significant differences in the data volume of each modality, which may restrict the performance of multi-modal large models on modalities with a small amount of data. Therefore, according to a possible design, when using various modality samples for comprehensive training, the balance of the number of samples can be achieved through a sample balancing strategy.
[0064] For example, in one embodiment, a part of the samples can be sampled from the samples with a larger number, so that the number of sampled samples is basically the same as the number of samples in the modality with a smaller number. For example, there are 1000 audio-text samples, 3 million text samples, and 100,000 image-text samples. Randomly sample about 1000 samples from the text samples and image-text samples respectively, and conduct comprehensive training of all modalities together with the audio-text samples.
[0065] In another embodiment, sample sampling can be performed according to the number of iteration steps. Specifically, in the parameter update cycle, single-modal samples can be sampled in a single batch, and the sampling batches of various modality samples are kept basically the same. For example, in a single parameter update cycle, the image-text samples, audio-text samples, and text samples are alternately sampled in the same batch (such as one batch), so that during the training process, the number of training batches of various modalities is stable and consistent. In a single parameter update cycle, the sampling quantity of a single batch of samples of various modalities can be the same, or the sampling probability can be positively correlated with the total sample data volume. In a specific example, in a single parameter update cycle, n text samples, n image-text samples, and n audio-text samples can be used, processed through a multi-modal large model in 3 batches, and the prediction results are obtained. The prediction results are compared with the corresponding label data to determine the model loss and adjust the model parameters. For another example, in a single parameter update cycle, the samples of one modality are used to determine the model loss and adjust the model parameters. The samples used in each parameter update cycle are rotated according to the modality. For example, n text samples are used in the first parameter update cycle, n image-text samples are used in the second parameter update cycle, n audio-text samples are used in the third parameter update cycle, n text samples are used in the fourth parameter update cycle... and so on.
[0066] In other embodiments, the balance of the number of samples of various modalities can also be achieved through other reasonable methods, which will not be elaborated here.
[0067] In an alternative embodiment, considering the differences between various modal samples, corresponding loss weights can also be set for the losses determined according to various modal samples to balance the processing ability differences of the large language model for various modal samples and enable the stable update of the multi-modal large model. When using multi-modal samples in a single parameter update cycle, the loss weights can also be used to weight the model losses in various modalities to determine the comprehensive model loss for adjusting the corresponding model parameters. Among them, the loss weights can be set in advance.
[0068] In one example, the loss weight of the text sample can be set to be the smallest, and the loss weight of the image-text sample can be set to be the largest.
[0069] In another example, the loss value ranges of various modalities can be used to determine the loss weights of the corresponding modal losses. The specific steps are as follows: The multi-modal large model is trained using a sample set of a predetermined scale in each modality until the corresponding convergence condition is met to obtain the corresponding model loss as the convergence loss in the corresponding modality, and the convergence loss values of each modality are recorded; for each modality, the normalized value of its convergence loss value relative to each convergence loss value is used as the loss weight in the corresponding modality. This loss weight determination process can be carried out before the start of the multi-modal large model training process or in the first several cycles.
[0070] For example, for a single modality i, a small sub-sample set is used to train the multi-modal model until it converges to meet the first convergence condition (such as the convergence of the multi-modal large model), and the model loss value at this time is recorded as the first convergence loss Then the corresponding normalized weight is calculated as the loss weight corresponding to modality i. For example, the normalized weight is calculated through the following formula: where M is the number of modalities, and α is a balance coefficient, such as set to 10.
[0071] During the parameter update process, the model losses determined under each modal sample can be multiplied by the loss weights and applied to the parameter update. For example, the update formula of the model parameter θ updated by the gradient descent method at the t-th step can be adjusted to:
[0072]
[0073] It should be noted that through the comprehensive training of various modal samples, the multi-modal large model can approach convergence. Through the phased and progressive training scheme from step 201 to step 203, the modal support ability of the multi-modal large model can be gradually expanded, and stable training of data in various modalities can be achieved.
[0074] In actual business operations, there may be a higher demand for processing precision. Therefore, in some possible designs, the training process of steps 201 to 203 described above can be used as the pre-training process of the multimodal large model, and the instruction fine-tuning of the multimodal large model can be carried out separately. In an alternative embodiment, when the training process of steps 201 to 203 is used as the pre-training process, the resolution of the image-text data can be set to a lower resolution (such as 256×512) to accelerate the training speed.
[0075] Continuing to refer to Figure 2 as shown, the steps 204 of instruction fine-tuning may include, for example: Step 2041, the image-text instruction fine-tuning step: using the image-text samples and text samples to train the visual encoder, the visual feature bridging module, and the large language model; Step 2042, the visual instruction fine-tuning step: using the image-text samples and video-text samples to train the visual encoder, the visual feature bridging module, and the large language model; Step 2043, the full-modal instruction fine-tuning step: using the image-text samples, audio-text samples, and text samples to train the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module, and the large language model.
[0076] Step 2041, Step 2042, and Step 2043 respectively correspond to Figure 4 Phase 1, Phase 2, and Phase 3 in Figure 3 Similarly, Figure 4 in the part with the flame pattern, the network whose model parameters are to be adjusted in the current phase is included, and in the part with the snowflake pattern, the network whose model parameters are considered fixed parameters in the current phase is included.
[0077] In step 2041, using the image-text samples and text samples, the model parameters in the visual encoder, the visual feature bridging module, and the large language model are adjusted.
[0078] Here, the text samples can be samples with the input data in pure text modality. After the multimodal large model is comprehensively trained with various modal samples, the multimodal large model has the corresponding processing ability for various modal data. At this time, further based on the image-text samples and text samples, the model parameters related to image-text data processing, such as the visual encoder and the visual feature bridging module, are fine-tuned, and the large language model is adaptively fine-tuned (as shown in Figure 4 Phase 1), which can enhance the processing ability of the multimodal large model for image-text data.
[0079] It should be noted that for the multimodal large model after pre-training, it already has the corresponding image-text processing ability. In this step 2041, image data with a higher resolution (such as 7680×4320) can be used to further fine-tune the model parameters and enhance the image-text processing ability.
[0080] In step 2042, using the image text samples and video text samples, adjust the model parameters in the visual encoder, visual feature bridging module, and large language model.
[0081] As Figure 4 shown in stage 2 of , in this step, the image text samples and video text samples can be used to perform instruction fine-tuning on the image-text processing network. Since a video can contain multiple images, the video can also be processed through the image-text processing network. The adjusted network still includes a visual encoder, a visual feature bridging module, and a large language model. The instruction fine-tuning in this stage can further enhance the image-text understanding ability of the multimodal large model.
[0082] In step 2043, using the image text samples, audio text samples, and text samples, adjust the model parameters in the visual encoder, visual feature bridging module, audio encoder, audio feature bridging module, and large language model.
[0083] After enhancing the image-text understanding ability of the multimodal large model through steps 2041 and 2042, further instruction fine-tuning of the entire network structure using various modal data can be performed to improve the processing ability of the multimodal large model for all modal data. Similar to step 203, the balance of the sample quantity can be achieved through the sample balancing strategy.
[0084] In addition, during the model convergence process, the model parameters of each network approach stable values. During the full-modal training process, the parameter updates of the processing networks for each modality usually depend on the samples under the corresponding modality, and the instruction fine-tuning of the large language model depends on the samples of various modalities. Therefore, there may be a situation where the convergence speeds of the model parameters in the processing networks corresponding to each modality are inconsistent. For this reason, this specification can also provide a dynamic equilibrium convergence control strategy to achieve the consistent convergence of model parameters during the full-modal training process.
[0085] This dynamic equilibrium convergence control strategy can control the consistency of model parameter convergence by adjusting the weights of the model losses corresponding to the samples under various modalities, thereby affecting the magnitude of the gradients for backpropagation. Specifically, the training for each modality can be regarded as a training task, and the principle of multi-task learning (MTL) can be used to balance the training progress for each modality (task), that is: for the modality showing a relatively slow convergence slope (smaller slope value), a lower training weight is given to prevent overfitting; for the modality showing a relatively steep convergence slope (larger slope value), a higher training weight is assigned to promote its learning.
[0086] Refer to Figure 5As shown, the samples of a single batch can be of a single modality, such as one of the image-text modality, text modality, audio modality, and video modality. Using samples of a single modality can determine a model loss. In a single parameter update cycle, samples of a single modality can be used to determine the current model loss, or samples of multiple modalities can be processed in multiple batches to determine the current model loss.
[0087] Figure 5 In the samples of the audio modality, the input audio can be converted into text data for processing, for example Figure 5 denoted as "Audio-Text" in []. For the samples of the text modality, the input data is Figure 5 denoted as "Text" in []. For the input data of the image-text modality samples, for example Figure 5 denoted as "Image-Text" in []. In the case of recognizing text from an image as the input data, it is Figure 5 denoted as "OCR" in []. For the input data in the video modality samples, for example Figure 5 denoted as "Video-Text" in []. And so on.
[0088] To use the current slope to regulate the training progress in the corresponding modality, this specification proposes a technical concept of determining the current slope corresponding to each modality through a validation set. Specifically, during the training process, periodic validation operations are inserted at fixed parameter adjustment cycle intervals. The interval between two validation operations is denoted as a validation cycle, and a single validation cycle can include multiple parameter adjustment cycles. In each validation cycle, a small validation subset is used to calculate the validation loss of various modalities, so that the training progress of various modalities can be tracked through the validation loss and convergence slope within the historical window. The validation subset can be a part of the sample data randomly segmented from the training data of the corresponding modality, including samples. Here, S i represents the number of validation batches in a single validation segment of the i-th modality, B i represents the validation batch size of this modality, and M represents the number of modalities. The data in the validation subset is excluded from the model training set.
[0089] Different modalities exhibit different degrees of training difficulty, resulting in different ranges of validation loss values. To ensure balanced initialization and reduce potential inaccuracies in the initial model loss curve, at the beginning of the training, specific loss weights such as w i,0 = 1 can be set for all modalities and kept at a fixed value in the first H validation segments.
[0090] At the beginning of the (H + 1)-th validation cycle, the weights of each modality can be determined through the slope. The formula a i,t x + b i,tA linear regression model of the form is used to fit the change in the validation loss within the historical window as the validation period \(t\) varies, and the slope coefficient \(a\) is obtained. i,t represents the current convergence rate of the modality. Based on this slope coefficient \(a\) i,t , the current slope of the loss curve for the corresponding modality in the current validation period can be determined. Figure 5 shows the case where the curve \((a\) i,t x + b\) i,t ) is fitted, where \(a\) i,t represents the slope of the fitted straight line, \(i = 1, 2, 3,\cdots,6\), corresponding respectively to Figure 5 the illustrated image - text modality, OCR, audio - text modality, text modality, interleaved - image, and video - text modality.
[0091] To ensure fair weight allocation across modalities, in one embodiment, the convergence slope of each modality can be calculated based on the normalized validation loss. The normalized validation loss for modality \(i\) is as follows:
[0092]
[0093] where represents the validation loss of the \(i\) - th modality at the \(t\) - th validation period, \(H\) represents the size of the historical window, that is, the number of historical validation periods referred to forward during the calculation of the validation loss for the current validation period, and \(\epsilon\) is a very small positive number, such as set to 10 #; to prevent division by zero. This normalized validation loss can be used to fit the change in the validation loss within the historical window as the validation period \(t\) varies through a linear regression model, and the obtained slope coefficient is denoted as \(a\) i,t .
[0094] To determine the current loss weight for each modality based on the current slope of the corresponding loss curve for each modality, the normalization results of the respective current slopes corresponding to each modality can be used to determine the respective convergence scores. A single convergence score is negatively correlated with the corresponding current slope. Then, the respective convergence scores are mapped to respective importance coefficients through an activation function, and the weight coefficient of the model loss for each modality is determined based on the product of the importance coefficient and the number of modalities. Next, the respective current loss weights corresponding to each modality within the current validation period are determined respectively according to the modality weights of each modality.
[0095] As a specific example, for the \(t\) - th validation period (where \(t>H\)), the normalized slope and the convergence score \(s\) i,t of the \(i\) - th modality can be calculated, for example, as follows:
[0096]
[0097] Among them, the softmax operation is performed on the modality dimension. Then, the modality weight assignment for the current validation period is calculated:
[0098] w i,t = M * softmaxBf * s i,t C
[0099] Among them, f is a scaling factor that adjusts the weight probability distribution. Multiplying by M ensures that the sum of the weights of all modalities is equal to M.
[0100] Adopting the method of the validation period can dynamically adjust the modality weights with minimal computational overhead, thereby improving the performance of all modalities in the context of full-modal learning.
[0101] In some alternative embodiments, in order to mitigate the sudden fluctuations caused by single-step weight updates and improve the stability of model training, an exponential moving average (EMA) mechanism can also be used to smoothly adjust the training weights of a single modality as follows:
[0102]
[0103] Among them, the smoothing factor α can be preset, such as 0.9. The adjusted modality loss weight is used for each training step in the next validation period.
[0104] In other embodiments, the slope corresponding to the loss curve can also be determined by other reasonable means, which will not be listed one by one here.
[0105] Through the forward processing of various modality samples by the multi-modal large model (M2-omini), the model loss can be determined. The product of the model loss and the corresponding current loss weight is used as the basis for determining the parameter update gradient in the backward process. Thus, the parameters in each neural network module in the multi-modal large model can be kept consistent during the regional convergence process, avoiding training biases caused by some networks converging quickly and some networks converging slowly.
[0106] Looking back on the above process, the training scheme of the multi-modal large model provided under the technical concept of this specification can decouple each network module in the multi-modal large model according to functions, and perform phased progressive training on the decoupled network modules, gradually expanding the modality support ability of the model and achieving better performance in each modality. This training method can effectively address the problem of large differences in data distribution among modalities and achieve stable training of modality data.
[0107] In a further embodiment, the technical problem of unbalanced amounts of training sample data for each modality can also be solved by the technical concept of balanced sampling according to the number of iteration steps. The technical problem of inconsistent convergence speeds for each modality during training can be solved by calculating the slope of the loss curve for each modality to measure the convergence speed of the modality and dynamically adjusting the loss weights for each modality according to the convergence speed.
[0108] According to an embodiment of another aspect, there is also provided a training device for a multi-modal large model. The device can be provided in a computer, terminal, or server with a certain computing capacity. The multi-modal large model therein can include: a visual encoder and a visual feature bridging module connected in sequence, an audio encoder and an audio feature bridging module connected in sequence, and a large language model connected to the visual feature bridging module and the audio feature bridging module. Figure 6 FIG. 600 shows a training device for a multi-modal large model according to an embodiment.
[0109] As Figure 6 shown, the device 600 may include:
[0110] An encoder alignment unit 601, configured to train the visual feature bridging module and the audio feature bridging module by using image-text samples and audio-text samples;
[0111] A graphic-text knowledge enhancement unit 602, configured to train the visual encoder, the visual feature bridging module, and the large language model by using image-text samples and text samples;
[0112] A full-modal joint training unit 603, configured to train the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module, and the large language model by using image-text samples, audio-text samples, and text samples.
[0113] According to a possible design, the device 600 may further include an instruction fine-tuning unit 604, configured to perform instruction fine-tuning on the multi-modal large model through the following steps:
[0114] Graphic-text instruction fine-tuning step: training the visual encoder, the visual feature bridging module, and the large language model by using image-text samples and text samples;
[0115] Visual instruction fine-tuning step: training the visual encoder, the visual feature bridging module, and the large language model by using image-text samples and video-text samples;
[0116] Full-modal instruction fine-tuning step: training the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module, and the large language model by using image-text samples, audio-text samples, and text samples.
[0117] It should be noted that Figure 6The device 600 shown corresponds to Figure 2 the method described, Figure 2 and the corresponding descriptions in the illustrated method embodiments also apply to the device 600 and will not be repeated here.
[0118] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having stored thereon a computer program, which, when executed on a computer, causes the computer to execute the method described in conjunction with Figure 2 etc.
[0119] According to an embodiment of yet another aspect, there is also provided a computing device including a memory and a processor, where the memory stores executable code, and when the processor executes the executable code, the method described in conjunction with Figure 2 etc. is implemented. Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0120] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification should be included within the protection scope of the technical concept of this specification.
Claims
1. A training method for a multimodal large model, the multimodal large model comprising: A visual encoder and a visual feature bridge module connected in sequence, an audio encoder and an audio feature bridge module connected in sequence, and a large language model connected to the visual feature bridge module and the audio feature bridge module; the method includes the following pre-training steps: Encoder alignment step: Use image text samples and audio text samples to train the visual feature bridge module and the audio feature bridge module; Image-text knowledge enhancement step: Use image-text samples and text samples to train the visual encoder, visual feature bridge module, and large language model; Full-modal joint training steps: Use image-text samples, audio-text samples, and text samples to train the visual encoder, visual feature bridging module, audio encoder, audio feature bridging module, and large language model.
2. The method of claim 1, wherein: After the pre-training step, the method further includes the following step of fine-tuning the multimodal large model: Fine-tuning steps for image-text instructions: Use image-text samples and text samples to train the visual encoder, visual feature bridge module, and large language model; Visual instruction fine-tuning step: Use image text samples and video text samples to train the visual encoder, visual feature bridge module and large language model; Full-modal instruction fine-tuning step: Use image-text samples, audio-text samples, and text samples to train the visual encoder, visual feature bridging module, audio encoder, audio feature bridging module, and large language model.
3. The method of claim 1, wherein: The all-modal joint training step includes: In a single parameter update cycle, the same batch is sampled in turn for image text samples, audio text samples, and text samples; Determine the model loss under the corresponding mode by processing a single modality sample in a single batch; The model losses under various modes are weighted to determine the comprehensive model loss, which is used to adjust the corresponding model parameters.
4. The method of claim 1, wherein: In the all-modal joint training step, the model loss determined for the single modality sample is multiplied and balanced with the pre-determined loss weight to determine the reverse transfer gradient of the model parameters; the loss weight corresponding to the model loss of the single modality is determined by the following method: When the multimodal large model is trained using a sample set of a predetermined size under the single modality until a first convergence condition is satisfied, a first convergence loss satisfying the first convergence condition is obtained; The loss weight corresponding to the model loss under the single mode is determined according to the ratio of the inverse of the first model loss to the sum of the inverses of the convergence losses under various modes.
5. The method of claim 2, wherein: The full-modal instruction fine-tuning step includes: During the model parameter adjustment process, the following dynamic equilibrium convergence control strategy is implemented: The current loss weight under the corresponding mode is determined according to the current slope of the corresponding loss curve under each mode, and the current loss weight under a single mode is negatively correlated with the corresponding current slope; The model loss corresponding to the samples in each mode is multiplied and balanced by the current loss weight and then used for model parameter adjustment.
6. The method of claim 5, wherein: The current slope of the loss curve for a single mode is determined by: Obtain the verification loss of the previous H verification cycles of the current verification cycle, where a single verification cycle includes multiple parameter update cycles under the single mode, and the verification loss of a single verification cycle is determined by the following normalization method: the model loss determined by the verification set for the single verification cycle minus the minimum value of the loss determined by the verification set for its historical H verification cycles, and the ratio of the maximum value to the minimum value of the model loss determined by the verification set for the historical H verification cycles; According to the slope coefficient of the verification loss in the verification cycle t determined by fitting the linear regression model, the current slope of the loss curve under the corresponding mode in the current verification cycle is determined.
7. The method according to claim 5 or 6, wherein: Determining the current loss weight under the corresponding mode according to the current slope of the corresponding loss curve under each mode includes: The normalized results of the current slopes corresponding to the respective modes are used to determine the corresponding convergence scores, and the single convergence score is negatively correlated with the corresponding current slope; Each convergence score is mapped to each importance coefficient through the activation function, and the weight coefficient of the model loss under various modes is determined according to the product of the importance coefficient and the number of modes; The current loss weights corresponding to each modality in the current verification cycle are determined according to the modal weights under various modalities.
8. The method of claim 7, wherein: Determining the respective current loss weights according to the weight coefficients under various modes includes: For a single mode, the current loss weight is determined using the exponential moving average method of the weight coefficient.
9. A training device for a multimodal large model, the multimodal large model comprising: A visual encoder and a visual feature bridge module connected in sequence, an audio encoder and an audio feature bridge module connected in sequence, and a large language model connected to the visual feature bridge module and the audio feature bridge module; the device comprises: An encoder alignment unit configured to train a visual feature bridge module and an audio feature bridge module using the image text sample and the audio text sample; A graphic knowledge enhancement unit configured to train a visual encoder, a visual feature bridge module, and a large language model using image text samples and text samples; The omnimodal joint training unit is configured to train the visual encoder, the visual feature bridging module, the audio encoder, the audio feature bridging module and the large language model using image-text samples, audio-text samples and text samples.
10. The device of claim 9, wherein: The device also includes an instruction fine-tuning unit configured to perform instruction fine-tuning on the multimodal large model through the following steps: Fine-tuning steps for image-text instructions: Use image-text samples and text samples to train the visual encoder, visual feature bridge module, and large language model; Visual instruction fine-tuning step: Use image text samples and video text samples to train the visual encoder, visual feature bridge module and large language model; Full-modal instruction fine-tuning step: Use image-text samples, audio-text samples, and text samples to train the visual encoder, visual feature bridging module, audio encoder, audio feature bridging module, and large language model.
11. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.
12. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Post-training method and device for multi-modal model
CN121145976A
Training method and apparatus for multi-modal large model
WO2026179596A1