Efficient hybrid parallel training method for multi-modal model

By conducting architectural analysis and computational load evaluation on multimodal models, and using a hybrid parallel training method, the problems of load imbalance, inefficiency and reduced accuracy in multimodal model training are solved, and an efficient and scalable training process is achieved.

CN120163267AActive Publication Date: 2025-06-17NANJING UNIV OF INFORMATION SCI & TECH +1

Patent Information

Application Number
CN202510647012.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

When the existing multimodal model training methods deal with heterogeneous computing loads, there are problems such as load imbalance, low training efficiency and reduced accuracy.

Method used

An efficient hybrid parallel training method is adopted to optimize the training process by performing architectural analysis and computational load evaluation of multimodal models, allocating computing resources and dividing pipeline stages, combining data parallelism and asynchronous pipeline parallelism strategies.

Benefits of technology

It significantly improves hardware resource utilization, improves training efficiency, reduces accuracy loss, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163267A_ABST
    Figure CN120163267A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient hybrid parallel training method for a multi-modal model, and belongs to the technical field of computers. The method comprises the steps that firstly, a multi-modal model architecture is divided into a plurality of independent modal sub-modules according to the multi-modal model architecture, and each sub-module processes one kind of modal data; secondly, formulating a computing power distribution scheme by statically analyzing the computing load of each sub-module; calculating time and parameter quantity of each layer of a modal sub-module are analyzed, then the layers are divided into a plurality of stages by using a partition algorithm, and asynchronous pipeline parallel and data parallel are combined to accelerate training; according to the method, the hardware resource utilization rate is remarkably improved through the calculation power distribution and hybrid parallel strategy for the modal calculation load, and the acceleration of the multi-modal model training process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-modal model training method, specifically an efficient hybrid parallel training method for multi-modal models, belonging to the field of computer technology. Background Art

[0002] Multi-modal models provide a comprehensive perspective for complex tasks by jointly learning data from different modalities, significantly enhancing the model's reasoning ability. However, multi-modal models consist of multiple sub-modules with different functions, and these sub-modules have significant differences in structure, computational requirements, and data processing methods. For example, some modality sub-modules use deep convolutional neural networks or Vision Transformers to process high-dimensional data, with a heavy computational load, while some modality sub-modules process serialized language information based on the Transformer architecture, with relatively lighter computational requirements. This heterogeneity poses many challenges to existing distributed training strategies.

[0003] Existing distributed training strategies mainly include data parallelism and pipeline parallelism. Data parallelism divides the training data into sub-batches and distributes them to different devices for parallel processing. Each device holds a complete model copy and updates the parameters through gradient synchronization. However, in multi-modal models, due to the computational load differences of different modalities, a simple gradient synchronization method causes some devices to wait too long, resulting in waste of computational resources. Pipeline parallelism divides the model into multiple stages and distributes them to different devices for processing, reducing the memory burden on a single device. However, in multi-modal models, the uneven computational load of sub-modules easily forms a performance bottleneck, and some stages are idle waiting for forward computation, reducing the training efficiency. In addition, although asynchronous pipeline parallelism can improve throughput, it may lead to a decrease in accuracy due to outdated gradients, especially in the initial stage of training, which has a more obvious impact.

[0004] Existing technologies have problems such as load imbalance, low training efficiency, and accuracy degradation in multi-modal model training. There is an urgent need for an efficient training method that can optimize for the heterogeneity of multi-modal models. Summary of the Invention

[0005] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide an efficient hybrid parallel training method for multi-modal models, which significantly improves the utilization rate of hardware resources through computing power allocation and hybrid parallel strategies for modal computational loads.

[0006] Technical Solution: An efficient hybrid parallel training method for multi-modal models of the present invention includes the following steps: Obtain a multi-modal model and a training dataset; Perform architecture analysis on the multi-modal model. According to the modal types of the input data, divide the multi-modal model into multiple independent modal sub-modules, where each modal sub-module is dedicated to processing a specific modal data; By statically analyzing the parameters of each modal sub-module, evaluate the computational load of each modal sub-module when processing global batch size data; Based on the results of the computational load analysis, formulate a computing power resource allocation plan to determine the number of devices required for each modal sub-module; According to the computing power resource allocation plan, adopt a partitioning algorithm for each modal sub-module to divide the modal sub-module into multiple consecutive pipeline stages; Divide the training process into a warm-up stage and an acceleration stage. Use the dataset to train the multi-modal model using data parallelism and asynchronous pipeline parallelism respectively.

[0007] Furthermore, the steps of formulating a computing power resource allocation plan based on the results of the computational load analysis to determine the number of devices required for each modal sub-module include: Adopt the proportional allocation principle. According to the computational load ratio of each modal sub-module and the total number of devices, allocate the corresponding number of devices.

[0008] Furthermore, the steps of adopting a partitioning algorithm for each modal sub-module according to the computing power resource allocation plan to divide the modal sub-module into multiple consecutive pipeline stages include: Take each modal sub-module independently as the input of the partitioning algorithm. Analyze the number of parameters of each layer of the modal sub-module. Use dynamic programming to determine the optimal partitioning boundary according to the results of the hierarchical analysis. Divide the modal sub-module into multiple consecutive pipeline stages according to the optimal partitioning boundary to ensure the computational load balance of each pipeline stage.

[0009] Furthermore, the steps of dividing the training process into a warm-up stage and an acceleration stage and using the dataset to train the multi-modal model using data parallelism and asynchronous pipeline parallelism respectively include: In the first 10% stage of the model training, adopt the data parallel strategy for training, use the linear warm-up learning rate, and when switching stages, the learning rate reaches the preset maximum value. This stage is recorded as the warm-up stage; The remaining training stage is used as the acceleration stage. Switch to the asynchronous pipeline parallel strategy. Allocate the multiple pipeline stages divided by the multi-modal model to their respective devices for asynchronous update, and at the same time allocate the training dataset to each modal sub-module.

[0010] Further, in the acceleration stage, each pipeline stage runs independently on its respective GPU and performs asynchronous forward calculation and backpropagation. Meanwhile, the learning rate enters the decay stage, and the cosine function is used to make the learning rate show a smooth decay curve, eventually decaying to close to 0. Gradient accumulation is introduced, and the calculation results of each pipeline stage are temporarily stored in the local buffer. When the gradients of a preset number of micro-batches are accumulated on the GPU in the pipeline, the gradient update is performed uniformly.

[0011] Advantageous effects: Compared with the prior art, the significant advantages of the present invention are as follows: In view of the heterogeneity of the multi-modal model, the present invention overcomes the deficiencies of the existing methods in load balancing, training efficiency, and accuracy maintenance through computing power allocation and hybrid parallel strategies. Specifically, by allocating computing power resources according to the computational load, separating modalities, and dividing pipeline stages, the present invention avoids the problems of excessive device waiting time and performance bottlenecks, thus significantly improving the utilization rate of hardware resources; by adopting asynchronous pipeline parallelism and gradient accumulation, the waiting time is reduced, the training process is accelerated, and the overall training time is shortened; at the same time, data parallelism and linear warm-up learning rate are adopted in the warm-up stage, and gradient accumulation and learning rate decay are introduced in the acceleration stage, effectively reducing the accuracy loss caused by asynchronous updates and ensuring the model performance. The hybrid parallel strategy adopted by the present invention means that the training process can be flexibly adjusted according to specific task requirements and hardware resources. Whether it is a small device or a large-scale cluster, this adaptability can ensure the efficient operation of the training process, and at the same time improve the scalability of the method, making it applicable to different-scale application scenarios. In summary, the present invention can significantly improve the utilization rate of hardware resources while accelerating the training process of the multi-modal model, ensuring the training accuracy, and also having excellent scalability, which is an important improvement to the existing distributed training strategies. Description of the Drawings

[0012] Figure 1 It is a flowchart of an efficient hybrid parallel training method for a multi-modal model; Figure 2 It is a schematic diagram of modality separation of a two-stream multi-modal model; Figure 3 It is a schematic diagram of using data parallelism for training in the warm-up stage; Figure 4 It is a schematic diagram of using asynchronous pipeline parallel training in the acceleration stage; Figure 5 It is a comparison chart of the iteration time of the method described in the present invention and the mainstream training strategies after normalizing the single iteration time of the method described in the present invention to 1; Figure 6 It is a comparison chart of the training accuracy of the method described in the present invention and asynchronous pipeline and synchronous pipeline training. Detailed Embodiments

[0013] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments.

[0014] An efficient hybrid parallel training method for a multimodal model described in this embodiment has a flowchart as Figure 1 shown, and this method includes the following steps: Step 1: Obtain the multimodal model and the training dataset.

[0015] For a given multimodal model, this multimodal model includes at least two of modalities such as text, speech, and image. Obtain the training dataset. If the training model is a three-stream multimodal model of speech-text-image, then it is necessary to obtain the datasets corresponding to text, speech, and image.

[0016] Step 2: Conduct an architecture analysis on the multimodal model. According to the modal type of the input data, divide the multimodal model into multiple independent modal sub-modules, and each modal sub-module is dedicated to processing a specific modal data.

[0017] Combined with Figure 2 shown, in an example, the modal type of the input data of the multimodal model can be any two or three of text, speech, and image. If there are two, they are respectively denoted as Modal 1 and Modal 2. Then the given multimodal model is divided into 2 independent modal sub-modules, denoted as Sub-module 1 and Sub-module 2, and each modal sub-module processes one modal data.

[0018] Step 3: Evaluate the computational load of each modal sub-module when processing global batch size data by statically analyzing the parameters of each modal sub-module.

[0019] In the example, by analyzing the model structure and parameters of each modal sub-module, calculate the number of floating-point operations (FLOPs) when processing a single sample, and use this as the main indicator of the computational load. According to the type of modal data processed, the two-stream multimodal large model used is divided into Modal Sub-module 1 and Modal Sub-module 2. For Modal Sub-module 1, such as VisionTransformer as the multimodal model, the FLOPs mainly come from the calculations of the self-attention mechanism and the feed-forward network, and can be estimated by parameters such as the number of attention heads, sequence length, and embedding dimension. For Modal Sub-module 2, such as Transformer, the FLOPs are concentrated on token embedding and attention calculations, and can be calculated according to the sequence length and number of layers. The calculation formulas for the number of floating-point operations of the convolutional layer, fully connected layer, and attention mechanism are respectively shown as follows: , In the formula, represents the number of floating-point operations of the convolutional layer, where and represent the height and width of the output feature map respectively. is the number of input channels. is the number of output channels. is the height and width of the convolutional kernel. The coefficient 2 means that each convolution operation includes one addition and one multiplication. The convolutional layer is the core part of the computation-intensive.

[0020] , In the formula, is the number of floating-point operations of the fully connected layer. In the Transformer model, the FLOPs calculation of the attention mechanism is relatively complex. is the number of input neurons. is the number of output neurons. The coefficient 2 also comes from multiplication and addition.

[0021] , In the formula, represents the number of floating-point operations of the attention mechanism. is the sequence length. is the model dimension.

[0022] According to the results of the computational load, that is, the floating-point operation counts of sub-module 1 and sub-module 2, a computing power resource allocation scheme is formulated. Usually, sub-modules with larger computational load and memory requirements will be given more computing resource support, while sub-modules with smaller load will be allocated fewer resources. This scheme provides a reasonable starting point for the training process.

[0023] Step 4: Based on the results of the computational load analysis, formulate a computing power resource allocation scheme to determine the number of devices required for each modal sub-module.

[0024] According to the results of the computational load, a computing power resource allocation scheme is formulated. Usually, sub-modules with larger computational load and memory requirements will be given more computing resource support, while sub-modules with smaller load will be allocated fewer resources. This scheme provides a reasonable starting point for the training process.

[0025] Furthermore, the steps of formulating a computing power resource allocation scheme based on the results of the computational load analysis to determine the number of devices required for each modal sub-module include: Adopt the proportional allocation principle, and allocate the corresponding number of devices according to the computational load ratio of each modal sub-module and the total number of devices.

[0026] In one example, if the computational load ratio of sub-module 1 to sub-module 2 is calculated as 2:3 and the total number of GPU devices is 20, then sub-module 1 is allocated 8 GPU devices and sub-module 2 is allocated 12 GPU devices.

[0027] Step 5: According to the computing power resource allocation scheme, use the partitioning algorithm for each modal sub-module to divide the modal sub-module into multiple consecutive pipeline stages.

[0028] Further, the steps of using the partitioning algorithm for each modal sub-module according to the computing power resource allocation scheme to divide the modal sub-module into multiple consecutive pipeline stages include: Take each modal sub-module independently as the input of the partitioning algorithm, analyze the number of parameters of each layer of the modal sub-module, and use dynamic programming to determine the optimal partitioning boundary according to the results of the hierarchical analysis. Assume that the modal sub-module has n layers and is divided into k pipeline stages. The computational load of each layer is obtained through hierarchical analysis. The formula is as follows: , In the formula, i represents the first i layers of the modal sub-module, starting from 1, indicating that dynamic programming considers sub-problems from 1 to the first i layers; j represents the number of divided stages, which cannot exceed the maximum allowed number of stages k and is at least 1; m represents that the last stage starts from the m-th layer until the i-th layer; is the computational load of the last stage, that is, from the m-th layer to the i-th layer; is the minimum bottleneck load when the first m - 1 layers are divided into j - 1 stages.

[0029] By enumerating all possible m, the partitioning algorithm selects the scheme that minimizes the bottleneck load of the current partition. Finally, gives the minimum bottleneck load when all n layers are divided into k stages. By backtracking the dp table, the specific partitioning boundary can be determined, and the modal sub-module is divided into multiple consecutive pipeline stages according to the optimal partitioning boundary to ensure the balance of the computational load of each pipeline stage.

[0030] In the example, k GPU devices are allocated to sub-module 1. Take sub-module 1 as the input of the partitioning algorithm, use the analysis algorithm to analyze the number of parameters of each layer of sub-module 1, and then use the partitioning algorithm to divide the modal sub-module into k consecutive stages with balanced computational load.

[0031] Based on the computing power allocation results, use the partitioning algorithm for each separated modality separately. At the same time, some lightweight modules may be combined together to ensure that each sub-modal module can obtain a suitable model partition and ensure that the computational load of each stage is as balanced as possible. Taking the dual-stream multi-modal model as an example, modalities with high computational load may be allocated separately, while modalities with low computational load are combined with other lightweight modules, thus avoiding a certain device from slowing down the overall training progress due to excessive load. After the sub-module is divided into consecutive stages, data transmission is only performed at the stage boundary, avoiding frequent global communication.

[0032] Step 6: Divide the training process into a warm-up stage and an acceleration stage. Use the dataset to train the multimodal model using data parallelism and asynchronous pipeline parallelism respectively.

[0033] In the early stage of training, due to the large span of parameter updates, the accuracy loss caused by asynchronous updates is more significant. As the model parameters gradually stabilize, the impact of asynchronous updates on accuracy will weaken. Therefore, in this example, the entire training process is divided into two stages: warm-up and acceleration, and different parallel configurations are used in different stages to optimize efficiency and accuracy.

[0034] Furthermore, the steps of dividing the training process into a warm-up stage and an acceleration stage, and using the dataset to train the multimodal model using data parallelism and asynchronous pipeline parallelism respectively include: Combine Figure 3 , in the first 10% stage of the model training. For example, if 200 training epochs are set, the first 20 epochs are the warm-up stage, and the data parallelism strategy is used for training. The linear warm-up learning rate is used, and at the stage switch, the learning rate reaches the preset maximum value. This stage is recorded as the warm-up stage; In the warm-up stage, the multimodal model is regarded as a whole and data parallelism is used to ensure the stability of parameter updates. In this stage, by reducing the impact of gradient obsolescence, the initial improvement of model accuracy is guaranteed.

[0035] Except for the warm-up stage, the remaining training stage is used as the acceleration stage, and the asynchronous pipeline parallelism strategy is switched. The multiple pipeline stages divided by the multimodal model are assigned to their respective devices for asynchronous updates, and at the same time, the training dataset is assigned to each modal sub-module.

[0036] When the model parameters gradually stabilize, switch to asynchronous pipeline parallelism and introduce the gradient accumulation technique to maximize the training throughput, thereby accelerating the overall training process.

[0037] Combining asynchronous pipeline parallelism and the stage training strategy can not only improve the training efficiency but also effectively alleviate the problem of accuracy decline caused by asynchronous pipeline parallelism. This method makes full use of hardware resources and provides an efficient and flexible solution for the distributed training of multimodal models.

[0038] Furthermore, in the acceleration stage, each pipeline stage runs independently on its own GPU and performs asynchronous forward calculation and backward propagation. At the same time, the learning rate enters the decay stage, and the learning rate shows a smooth decay curve through the cosine function, and finally smoothly decays to close to 0. The gradient accumulation is introduced, and the calculation results of each pipeline stage are temporarily stored in the local buffer. When the gradients of the preset number of micro-batches are accumulated on the GPU in the pipeline, the gradient update is performed uniformly.

[0039] After completing the model partitioning, an asynchronous pipeline parallel strategy is introduced during the acceleration phase to speed up the training process. This strategy allows each computing device to process tasks at different stages without global synchronization, thus maximizing device utilization and reducing idle time. The training data batches are split into multiple micro-batches and injected into the pipeline sequentially. When processing the current micro-batch, each device can asynchronously pass the output of the previous micro-batch to the downstream device while starting to process the next micro-batch. This effectively overlaps the computing and communication times and hides the communication latency. During backpropagation, the gradient calculation and parameter update are performed asynchronously. The device updates the local parameters immediately after completing the gradient calculation of a micro-batch without waiting for other devices to synchronize, further reducing the synchronization waiting time. Through this asynchronous mechanism, this strategy can achieve efficient multi-device collaboration in both forward and backward propagations, make full use of computing resources, and significantly accelerate the overall training.

[0040] After the warm-up phase ends, switch to the acceleration phase, as Figure 4 Taking the dual-stream multi-modal model as an example, according to the allocation results of each modal sub-module, sub-module 1 and sub-module 2 are divided into multiple pipeline stages, asynchronous pipeline parallelism is used, and gradient accumulation technology is introduced to mitigate the impact of weight obsolescence on accuracy loss. During the asynchronous pipeline stage, a smaller learning rate is used, and the expression is: , where lr is the learning rate for this round, is the preset learning rate, , and are the current step number, switching step number, and total step number respectively. The learning rate shows a smooth decay curve through the cosine function and finally decays smoothly to close to 0. A smaller learning rate can reduce the impact of weight obsolescence caused by too large a parameter update amplitude.

[0041] To further illustrate the effectiveness and superiority of the efficient hybrid parallel training method for multi-modal models described in the present invention, the following examples are used for illustration. The multi-modal models used are the CLIP and LIT multi-modal models respectively, and the preprocessed food101 dataset is used. The training set contains 75,750 pictures and the experiment is carried out on 8 H800 GPUs of a 1-machine 8-card server. Comparing with the existing 1F1B and Gpipe pipeline parallel training methods, the iteration speed of a single round has been improved, and the relevant data in the model configuration table are shown in Table 1 below.

[0042] Table 1

[0043] Combined with Figure 5 and Figure 6, using CLIP with ResNet101 as image mode and Transformer as text mode, training 60 rounds on the dataset Country211, with learning rate set to 5e-3, momentum 0.9, weight decay 1e-4, and batch size . Figure 5 This is an iteration time comparison chart after the single iteration time of the method is standardized to 1. From the figure, it can be seen that the method of the present invention (denoted as HybPipe) is improved compared with the mainstream training strategies, especially in the two-node environment. Taking the LIT model as an example, in the single-node environment, that is, Nodes=1, the single iteration time of 1F1B and Gpipe is 1.083 times and 1.125 times slower than that of the present method, respectively. In the multi-node environment, that is, Nodes=2, the iteration time gap is further widened, and the single iteration time of 1F1F and Gpipe is 1.182 times and 1.257 times slower than that of the method of the present invention, respectively. Figure 6 In the figure, within the same iteration rounds, i.e., 60 rounds, the final accuracy rates of 1F1B, asynchronous pipeline Asynchronous and HybPipe of the present invention are 73.81%, 69.71% and 74.4% respectively. It can be seen that the accuracy rate of the model trained by the method of the present invention is consistent with the accuracy rate of the synchronous pipeline strategy training, thus avoiding the problem of decreased accuracy caused by gradient obsolescence caused by asynchronous pipeline parallelism.

[0044] The present invention proposes an innovative training method, which cleverly combines asynchronous pipeline parallel technology and staged training strategy. Through this combination, not only the training efficiency is significantly improved, but also the problem of decreased model accuracy that may be caused by asynchronous pipeline parallelism is successfully alleviated. In addition, the method can make full use of hardware resources and provide an efficient and flexible solution for distributed training of multimodal models. When training a dual-stream multimodal model on two devices with 1 machine and 8 cards, the entire model is distributed to all GPU devices using 16 GPU data parallelism in the warm-up stage; in the acceleration stage, each modal sub-module of the model is divided into multiple consecutive stages, and each stage independently occupies a GPU using asynchronous pipeline parallelism. By injecting micro-batches into the pipeline, the overlap of GPU device calculation and communication is achieved, the idle time of the GPU device is reduced, and resource utilization is improved.

[0045] The core of the multimodal model is to combine data from different modalities to achieve joint learning and reasoning. By integrating data from multiple modalities into the same model, the model has a more comprehensive perspective when dealing with complex tasks. However, there are significant differences in characteristics, dimensions, and structures among data from different modalities, and there are many limitations in existing distributed model training strategies. The present invention provides an efficient hybrid parallel training strategy for multimodal models, which optimizes the heterogeneous characteristics of multimodal models by combining data parallelism, asynchronous pipeline parallelism, and staged training strategies, and realizes the acceleration of the training process of multimodal models. The present invention can be applied to the following fields: Acceleration of multimodal model training: During the training process of multimodal models, an efficient hybrid parallel strategy can significantly accelerate the training process and reduce time and resource consumption under large-scale datasets and complex model architectures. This technology is particularly suitable for industrial-level artificial intelligence applications that require rapid iteration and deployment.

[0046] Cross-modal data analysis: In scenarios where multiple data types need to be processed simultaneously, the cross-modal reasoning ability of the model can be improved through efficient parallel training. For example, in audio content analysis, the present invention can synchronously process text (text modality) and audio tracks (auditory modality), and generate comprehensive understanding results by combining text descriptions. This ability is of great significance for applications such as intelligent monitoring and multimedia content recommendation.

Claims

1. An efficient hybrid parallel training method for multimodal models, characterized in that: The steps include: Obtain multimodal models and data sets required for training; Performing architecture analysis on the multimodal model, dividing the multimodal model into multiple independent modal submodules according to the modal type of the input data, each modal submodule is dedicated to processing a specific modal data; By statically analyzing the parameters of each modality submodule, the computational load of each modality submodule when processing global batch size data is evaluated; Based on the results of the computational load analysis, a computing resource allocation plan is developed to determine the number of devices required for each modal submodule; According to the computing power resource allocation plan, a partitioning algorithm is used for each modal submodule to divide the modal submodule into multiple continuous pipeline stages; The training process is divided into a warm-up phase and an acceleration phase. The multimodal model is trained using data parallelism and asynchronous pipeline parallelism using the dataset.

2. The efficient hybrid parallel training method for a multimodal model according to claim 1, characterized in that: Based on the results of the computational load analysis, the steps to formulate a computing resource allocation plan and determine the number of devices required for each modal submodule include: The proportional allocation principle is adopted to allocate the corresponding number of devices according to the calculation load ratio of each modal sub-module and the total number of devices.

3. The efficient hybrid parallel training method for a multimodal model according to claim 1, characterized in that: According to the computing power resource allocation scheme, a partitioning algorithm is used for each modal submodule, and the steps of dividing the modal submodule into multiple consecutive pipeline stages include: Each modal submodule is used as the input of the partitioning algorithm independently, and the parameter size of each layer of the modal submodule is analyzed. Dynamic programming is used to determine the optimal partition boundary based on the results of the hierarchical analysis. The modal submodule is divided into multiple continuous pipeline stages based on the optimal partition boundary.

4. The efficient hybrid parallel training method for a multimodal model according to claim 1, characterized in that: The training process is divided into a warm-up phase and an acceleration phase. The steps of using the data set to train the multimodal model using data parallelism and asynchronous pipeline parallelism include: In the first 10% of the model training, the data parallel strategy is used for training, and the linear warm-up learning rate is used. When the stage is switched, the learning rate reaches the preset maximum value. This stage is recorded as the warm-up stage; The remaining training phase is used as an acceleration phase, switching to an asynchronous pipeline parallel strategy. Multiple pipeline stages divided by the multimodal model are allocated to their respective devices for asynchronous updates, and the training dataset is allocated to each modality sub-module.

5. An efficient hybrid parallel training method for a multimodal model according to any one of claims 1 to 4, characterized in that: In the acceleration phase, each pipeline stage runs independently on its own GPU and performs asynchronous forward calculation and back propagation. At the same time, the learning rate enters the decay phase. The cosine function is used to make the learning rate present a smooth decay curve, and finally decays to close to 0. Gradient accumulation is introduced, and the calculation results of each pipeline stage are temporarily stored in the local buffer. When the pipeline on the GPU accumulates the gradient of a preset number of micro-batches, the gradient is updated uniformly.

Citation Information

Patent Citations

  • Hybrid pipeline parallel method for accelerating distributed deep neural network training

    CN112784968A

  • Distributed deep learning method and device based on multiple GPUs and electronic equipment

    CN114820279A

  • Multi-modal data analysis method based on hybrid expert structure large model training

    CN118551220A

  • Machine learning systems and methods for deep learning of genomic contexts

    US20240312558A1

Cited By

  • Distributed training method of multi-modal model, electronic equipment and storage medium

    CN121144857A

  • Heterogeneous end side equipment-oriented multi-modal model assembly line parallel training method

    CN122286313A

  • A multi-modal model pipeline parallel training method for heterogeneous end-side devices

    CN122286313B