Large model distributed separation training method and device
By dividing the large language model into a sequence of stages and replicating the video encoder in each stage, independently processing image data and transferring gradients, the problem of low training efficiency of large multimodal models is solved, and efficient resource allocation and computational optimization are achieved.
Patent Information
- Application Number
- CN202510805445.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The training efficiency of large multimodal models is low, especially because the deep structure of the visual encoder leads to computational bottlenecks during pipeline parallel training.
The large language model is divided into a sequence of stages, and the video encoder is replicated at each stage. The image data is independently processed to generate image features, and training is performed through gradient transfer, separating the training process of the large language model and the video encoder.
It significantly improves the overall training efficiency of large multimodal models, eliminates computing bottlenecks, and improves resource utilization and training throughput.
Smart Images

Figure CN120599441A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the field of multimodal large model technology. Background Art
[0002] Building a universal model that can effectively perceive the world through multimodal signals has been a long-standing goal. The architecture of a large multimodal model for late-stage fusion connects a visual encoder to a large language model. The video encoder processes image data to generate image features, which are then concatenated with text features and fed into the large language model for autoregressive training.
[0003] Large multimodal models have a huge number of parameters. To improve model training efficiency, a distributed architecture called PipelineParallel (PP) is often used to train these models. Because the visual encoder has far fewer parameters than a large language model, it is often placed alongside the embedding layer of the large language model in PipelineParallel (PP0). However, the depth of the visual encoder can make PP0 a computational bottleneck. Summary of the Invention
[0004] The embodiments of the present disclosure provide a large-model distributed separation training method, apparatus, device, storage medium, and program product.
[0005] In a first aspect, an embodiment of the present disclosure proposes a large model distributed separation training method, comprising: dividing a large language model into a stage sequence, and replicating a video encoder at each stage in the stage sequence; inputting image data into the video encoder in the stage sequence to generate image features; using the image features to train the large language model part in the stage sequence, and transmitting the gradient of the image features back to the video encoder in the stage sequence during the training process; and using the gradient of the image data and the image features to train the video encoder in the stage sequence.
[0006] In the second aspect, an embodiment of the present disclosure proposes a large model distributed separation training device, including: a segmentation module, configured to segment the large language model into a stage sequence, and replicate the video encoder at each stage in the stage sequence; an encoding module, configured to input image data into the video encoder in the stage sequence to generate image features; a first training module, configured to use the image features to train the large language model part in the stage sequence, and pass the gradient of the image features back to the video encoder in the stage sequence during the training process; a second training module, configured to use the gradient of the image data and the image features to train the video encoder in the stage sequence.
[0007] In a third aspect, an embodiment of the present disclosure proposes an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect.
[0008] In a fourth aspect, an embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute the method described in the first aspect.
[0009] In a fifth aspect, an embodiment of the present disclosure proposes a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.
[0010] The key or important features of the embodiments of the present disclosure are not intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Other features, objects, and advantages of the present disclosure will become more apparent upon reading the detailed description of the non-limiting embodiments made with reference to the following drawings. The drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. Among them: Figure 1 is a flowchart of an embodiment of a large model distributed separation training method according to the present disclosure; Figure 2 is a flowchart of another embodiment of the large model distributed separation training method according to the present disclosure; Figure 3 This is a flowchart of the distributed separation training of large models; Figure 4 is a structural diagram of an embodiment of a large-scale model distributed separation training device according to the present disclosure; Figure 5 It is a block diagram of an electronic device used to implement the large-model distributed separation training method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0012] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0013] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0014] Figure 1 A process 100 of an embodiment of a large model distributed separation training method according to the present disclosure is shown. The large model distributed separation training method includes the following steps: In step 101 , the large language model is divided into a stage sequence, and a video encoder is replicated at each stage in the stage sequence.
[0015] In this embodiment, the execution entity of the large model distributed separation training method can divide the large language model into a stage sequence and replicate the video encoder at each stage in the stage sequence.
[0016] A large multimodal model can consist of a visual encoder and a large language model. Its architecture can be to connect the visual encoder to the large language model. The large multimodal language model can process both image and text data. The visual encoder can be used to process images, while the large language model is used to process text.
[0017] The large language model is sequentially split into multiple parts, each of which is sequentially copied into a stage, generating a stage sequence (PP0, PP1, PP2, ..., PPn). Unlike existing techniques where the visual encoder is placed in PP0, the video encoder is copied to each stage in the stage sequence. In other words, a complete copy of the video encoder is made at each stage in the stage sequence, with the same initial parameters.
[0018] Step 102: input the image data into a video encoder in a sequence of stages to generate image features.
[0019] In this embodiment, the execution entity may input image data into a video encoder in a sequence of stages to generate image features, wherein the video encoder in each stage may be used to process a portion of the image data.
[0020] In order to ensure that the video encoder in each stage of the stage sequence processes a portion of the data, the image data can be segmented along the stage dimension to generate sub-image data corresponding to each stage. For example, if the stage sequence includes four stages, the image data can be segmented into four parts, with each stage processing one part. The sub-image data corresponding to each stage is then forward-propagated through the video encoder in the corresponding stage to generate sub-image features corresponding to each stage, and the sub-image features corresponding to each stage are merged to generate image features. The video encoder in each stage processes the corresponding sub-image data independently, but the parameters of the video encoder stop gradient and are not trained.
[0021] Step 103: Use the image features to train the large language model part in the stage sequence, and during the training process, the gradient of the image features is transmitted back to the video encoder in the stage sequence.
[0022] In this embodiment, the execution entity may utilize image features to train the large language model portion in the stage sequence, and transmit the gradient of the image features back to the video encoder in the stage sequence during the training process.
[0023] Based on image features, input data for the large language model can be generated. By forward-propagating the input data through the large language model in the stage sequence, the gradient of the large language model can be generated. By back-propagating the gradient of the large language model through the large language model in the stage sequence, the parameters of the large language model can be updated. Furthermore, during the training of the large language model, the gradient of the image features can be passed back to the video encoder in the stage sequence.
[0024] During the training of the large language model, the parameters of the video encoder are set to stop gradients. This means that the gradients of the large language model are not automatically backpropagated to the video encoder. The video encoder is only responsible for providing image features during this forward propagation, and its parameters are temporarily frozen.
[0025] Step 104: Use the image data to train the video encoder in the stage sequence.
[0026] In this embodiment, the execution entity may use image data to train the video encoder in the stage sequence.
[0027] The video encoder in the sequence of stages reprocesses the image data. The image data is forward propagated through the video encoder in the sequence of stages to generate image features again. The video encoder in the sequence of stages then uses the gradients of the returned image features to manually perform backpropagation through the video encoder and update the video encoder parameters.
[0028] The disclosed embodiments provide a large-scale distributed separate training method that separates large language model training from video encoder training and employs a differentiated distributed parallel strategy to optimize resource allocation. Video encoder training no longer blocks the large language model's pipeline parallelization process, allowing large language model training to maintain the efficient standard pipeline parallelization mode, significantly improving the overall training efficiency of large multimodal models.
[0029] Continue to refer Figure 2 , which shows a process 200 of another embodiment of the large model distributed separation training method according to the present disclosure. The large model distributed separation training method includes the following steps: In step 201 , the large language model is divided into a stage sequence by layer using pipeline parallelism, and a video encoder is replicated at each stage in the stage sequence.
[0030] In this embodiment, the execution body of the large model distributed separation training method can use pipeline parallelism to divide the large language model into a stage sequence by layer, and replicate the video encoder at each stage in the stage sequence.
[0031] A large multimodal model can consist of a visual encoder and a large language model. Its architecture can be to connect the visual encoder to the large language model. The large multimodal language model can process both image and text data. The visual encoder can be used to process images, while the large language model is used to process text.
[0032] Large language models can include multiple layers. Using PP technology, large language models can be divided into multiple parts (called stages) in order of layers, generating a sequence of stages (PP0, PP1, PP2, ..., PPn). Each stage can include one or more layers of the large language model and is deployed on a graphics processing unit (GPU). Data flows between GPUs in a pipelined manner, with each GPU processing a stage of the model and then passing the results to the next GPU.
[0033] Unlike existing techniques where the visual encoder is placed in PP0, the video encoder is replicated at each stage in the sequence. In other words, a complete copy of the video encoder is replicated at each stage in the sequence, with the same initial parameters.
[0034] Step 202: segment the image data based on the stage dimension to generate sub-image data corresponding to each stage.
[0035] In this embodiment, the execution entity may segment the image data based on the stage dimension to generate sub-image data corresponding to each stage.
[0036] To ensure that the video encoder in each stage of the stage sequence processes a portion of the data, the image data can be split along the stage dimension to generate sub-image data corresponding to each stage. For example, if the stage sequence includes four stages, the image data can be split into four parts, with each stage processing one part.
[0037] Step 203 : forward-propagate the sub-image data corresponding to each stage in the video encoder in the corresponding stage to generate sub-image features corresponding to each stage.
[0038] In this embodiment, the execution entity may forward-propagate the sub-image data corresponding to each stage through the video encoder in the corresponding stage to generate sub-image features corresponding to each stage. The video encoder in each stage independently processes the corresponding sub-image data, but the parameters of the video encoder are de-graded and not trained.
[0039] Step 204 : Aggregate the sub-image features corresponding to each stage into the first stage in the stage sequence to generate image features.
[0040] In this embodiment, the execution entity may aggregate the sub-image features corresponding to each stage into the first stage PP0 in the stage sequence (PP0, PP1, PP2, ..., PPn) to generate image features.
[0041] Step 205: Generate input data by combining image features and text features.
[0042] In this embodiment, the execution entity may generate input data based on splicing of image features and text features.
[0043] Usually, image features and text features have different dimensions. In order to be able to splice image features with text features, the image features can be input into a resampler for resampling to obtain resampled image features. The resampler can be used to align image features with text features. Then, the resampled image features and text features are spliced in the sequence dimension to generate input data. For example, image features After the resampler, the hidden layer size of the video encoder can be aligned Hidden layer size of the embedding layer of the large language model , get the resampled image features . Resample image features With text features Splice in the sequence dimension to generate input data .
[0044] In step 206, the input data is forward propagated in the large language model portion in the stage sequence to generate the gradient of the large language model.
[0045] In this embodiment, the execution entity may forward propagate the input data in the large language model portion in the stage sequence to generate a gradient of the large language model.
[0046] The input data can flow forward through PP0, PP1, PP2, ..., PPn in sequence to generate the gradient of the large language model.
[0047] In step 207 , the gradient of the large language model is back-propagated in the large language model portion in the stage sequence to update the parameters of the large language model.
[0048] In this embodiment, the execution entity may back-propagate the gradient of the large language model in the large language model portion in the stage sequence to update the parameters of the large language model.
[0049] The gradient of the large language model can flow backward through PPn, ..., PP2, PP1, PP0 in sequence to update the parameters of the large language model.
[0050] During the training of the large language model, the parameters of the video encoder are set to stop gradients. This means that the gradients of the large language model are not automatically backpropagated to the video encoder. The video encoder is only responsible for providing image features during this forward propagation, and its parameters are temporarily frozen.
[0051] In step 208 , the gradient of the large language model is back-propagated to the first stage in the stage sequence to generate the gradient of the image features.
[0052] In this embodiment, the execution entity may back-propagate the gradient of the large language model to the first stage PP0 in the stage sequence to generate the gradient of the image feature.
[0053] During the training of the large language model, the gradient of the large language model is back-propagated from PPn to PP0. When the back-propagation reaches PP0, the gradient of the image feature can be obtained.
[0054] Step 209 : Back-propagate the gradient of the image features among the video encoders in the stage sequence.
[0055] In this embodiment, the execution entity may back-propagate the gradient of the image features between the video encoders in the stage sequence.
[0056] The gradient of the image feature on PP0 can be broadcast back between PPs, returning the gradient of the image feature to the video encoder on each PP.
[0057] In step 210 , the image data is forward propagated in the video encoder in the stage sequence, the gradient of the image features is manually back-propagated in the video encoder in the stage sequence, and the parameters of the video encoder in the stage sequence are updated.
[0058] In this embodiment, the execution entity may forward propagate the image data in the video encoder in the stage sequence, manually backpropagate the gradient of the image features in the video encoder in the stage sequence, and update the parameters of the video encoder in the stage sequence.
[0059] The video encoder in each stage reprocesses the corresponding sub-image data. The sub-image data corresponding to each stage is forward-propagated through the video encoder in that stage, regenerating the sub-image features corresponding to each stage. The video encoder in each stage then manually performs backpropagation of the video encoder using the gradients of the returned sub-image features. Through manual backpropagation, the gradients of the video encoder in each stage are calculated, and the parameters of the video encoder in that stage are independently updated based on the gradients of the video encoder in each stage. Finally, the video encoder parameters are synchronized between stages to complete the training of the video encoder.
[0060] The training of the video encoder is essentially a data-parallel training method. The video encoder in each stage uses the gradient of the corresponding sub-image features to update the local model copy, and the training of the video encoder is completed by synchronizing the parameters of the video encoder between stages.
[0061] The disclosed embodiments provide a large-model distributed separation training method that separates large language model training from video encoder training and adopts a differentiated distributed parallel strategy to optimize resource allocation. The training of the video encoder no longer blocks the pipeline parallel process of the large language model. The training of the large language model can maintain an efficient standard pipeline parallel mode, thereby significantly improving the overall training efficiency of the multimodal large language model. The video encoder is trained in parallel on a sequence of stages, utilizing the computing resources of all devices to process its deeper layers. The video encoder in each stage processes a portion of the image data, achieving data parallel processing and improving the processing efficiency of the image data. Since the computational load of the video encoder is distributed to all stages for parallel processing, the problem of PPO becoming a bottleneck due to the long computation time of the video encoder in traditional solutions is completely eliminated, significantly reducing cavitation during the pipeline training process. During the video encoder training stage, all stages are performing calculations simultaneously, which can fully utilize the computing density of the GPU and greatly improve the resource utilization and training throughput of the overall cluster.
[0062] Figure 3 The flowchart of the distributed separation training of large models is shown in FIG. Figure 3 As shown in the figure, a heterogeneous distribution strategy is designed. The LLM (Large Language Model) part is normally split into PP architectures. The VIT (Vision Transformer, a transformer-based visual encoder) is replicated on each PP.
[0063] In the first fwd (forward), the image data will be split on PP and img_fea (image features) will be obtained through VIT. Then all img_fea will be gathered to PP0 to obtain Acc img_fea.
[0064] In Acc When fwd, Acc img_fea flows forward through PP0, PP1, PP2, ..., PPn in sequence, generating Acc bwd. Then Acc bwd flows in reverse through PPn, ..., PP2, PP1, and PP0 in turn, updating the LLM parameters. It should be noted that during the training phase of the LLM backbone network, the VIT parameters will stop_gradient (stop gradient) and no training will be performed.
[0065] During the LLM backbone network training process, Acc bwd back propagates to PP0, and img_fea_grad can be obtained , and then reverse broadcast between PPs to convert img_fea_grad Return the VIT on each PP.
[0066] In the second fwd, each PP re-forwards the image data separately, and then uses the returned img_fea_grad Perform manual backpropagation to calculate the gradient of VIT parameters and update the parameters of VIT so that VIT can be trained in parallel among PPs.
[0067] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a large model distributed separation training device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0068] like Figure 4 As shown, the large model distributed separation training device 400 of this embodiment may include: a segmentation module 401, an encoding module 402, a first training module 403, and a second training module 404. The segmentation module 401 is configured to segment the large language model into a stage sequence and replicate the video encoder at each stage in the stage sequence; the encoding module 402 is configured to input image data into the video encoder in the stage sequence to generate image features; the first training module 403 is configured to use the image features to train the large language model portion in the stage sequence and transmit the gradient of the image features back to the video encoder in the stage sequence during the training process; the second training module 404 is configured to use the image data and the gradient of the image features to train the video encoder in the stage sequence.
[0069] In this embodiment, in the large model distributed separation training device 400, the specific processing of the segmentation module 401, the encoding module 402, the first training module 403 and the second training module 404 and the technical effects thereof can be referred to respectively. Figure 1 The relevant descriptions of steps 101-104 in the corresponding embodiment are not repeated here.
[0070] In some optional implementations of this embodiment, the segmentation module 401 is further configured to: utilize pipeline parallelism to segment the large language model into a stage sequence by layer, wherein each stage includes at least one layer of the large language model.
[0071] In some optional implementations of this embodiment, the encoding module 402 is further configured to: divide the image data in the stage dimension to generate sub-image data corresponding to each stage; forward propagate the sub-image data corresponding to each stage in the video encoder in the corresponding stage to generate sub-image features corresponding to each stage; and merge the sub-image features corresponding to each stage to generate image features.
[0072] In some optional implementations of this embodiment, the encoding module 402 is further configured to aggregate sub-image features corresponding to each stage into the first stage in the stage sequence to generate image features.
[0073] In some optional implementations of this embodiment, the first training module 403 is further configured to: generate input data based on image features; forward propagate the input data in the large language model part in the stage sequence to generate the gradient of the large language model; and backpropagate the gradient of the large language model in the large language model part in the stage sequence to update the parameters of the large language model.
[0074] In some optional implementations of this embodiment, the first training module 403 is further configured to: generate input data based on splicing of image features and text features.
[0075] In some optional implementations of this embodiment, the first training module 403 is further configured to: input the image features into the resampler for resampling to obtain resampled image features, wherein the resampler is used to align the image features with the text features; and splice the resampled image features and the text features in the sequence dimension to generate input data.
[0076] In some optional implementations of this embodiment, the first training module 403 is further configured to: back-propagate the gradient of the large language model to the first stage in the stage sequence to generate the gradient of the image features; and back-propagate the gradient of the image features between the video encoders in the stage sequence.
[0077] In some optional implementations of this embodiment, the second training module 404 is further configured to: forward propagate the image data in the video encoder in the stage sequence, manually backpropagate the gradient of the image features in the video encoder in the stage sequence, and update the parameters of the video encoder in the stage sequence.
[0078] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0079] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0080] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0081] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. Computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to bus 504.
[0082] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0083] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the large-model distributed separate training method. For example, in some embodiments, the large-model distributed separate training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the large-model distributed separate training method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the large-model distributed separate training method by any other suitable means (e.g., via firmware).
[0084] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0085] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0086] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0087] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0088] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0089] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0090] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved. This is not limited herein.
[0091] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A large model distributed separation training method, comprising: Split the large language model into a sequence of stages and replicate the video encoder at each stage in the sequence; Inputting image data into a video encoder in the sequence of stages to generate image features; Using the image features, training a large language model portion in the stage sequence, and transmitting the gradient of the image features back to the video encoder in the stage sequence during the training process; A video encoder in the sequence of stages is trained using the image data and gradients of the image features.
2. The method according to claim 1, wherein The large language model is divided into a sequence of stages, including: By utilizing pipeline parallelism, the large language model is divided into the stage sequence by layer, wherein each stage includes at least one layer of the large language model.
3. The method according to claim 1, wherein The step of inputting image data into a video encoder in the sequence of stages to generate image features comprises: Segmenting the image data in a stage dimension to generate sub-image data corresponding to each stage; Forward propagating the sub-image data corresponding to each stage in the video encoder in the corresponding stage to generate sub-image features corresponding to each stage; The sub-image features corresponding to each stage are merged to generate the image features.
4. The method according to claim 3, wherein: The merging of sub-image features corresponding to each stage to generate the image features includes: The sub-image features corresponding to each stage are aggregated into the first stage in the stage sequence to generate the image features.
5. The method according to claim 1, wherein The step of training the large language model portion in the stage sequence by utilizing the image features includes: generating input data based on the image features; Forward propagating the input data through the large language model portion in the stage sequence to generate a gradient of the large language model; Back-propagating the gradient of the large language model in the large language model part in the stage sequence to update the parameters of the large language model.
6. The method according to claim 5, wherein: Generating input data based on the image features includes: The input data is generated by splicing the image features and the text features.
7. The method according to claim 6, wherein: The step of generating the input data by combining the image features and the text features includes: Inputting the image features into a resampler for resampling to obtain resampled image features, wherein the resampler is used to align the image features with the text features; The resampled image features and the text features are concatenated in a sequence dimension to generate the input data.
8. The method according to claim 1, wherein The step of transmitting the gradient of the image feature back to the video encoder in the sequence of stages during training comprises: Back-propagating the gradient of the large language model to the first stage in the stage sequence to generate the gradient of the image feature; The gradients of the image features are back-propagated through the video encoders in the sequence of stages.
9. The method according to claim 1, wherein The step of training the video encoder in the sequence of stages using the image data and the gradient of the image features comprises: The image data is forward propagated through the video encoder in the stage sequence, the gradient of the image feature is manually back-propagated through the video encoder in the stage sequence, and the parameters of the video encoder in the stage sequence are updated.
10. A large-model distributed separation training device, comprising: a segmentation module configured to segment the large language model into a sequence of stages and replicate the video encoder at each stage in the sequence of stages; an encoding module configured to input image data into a video encoder in the sequence of stages to generate image features; A first training module is configured to train a large language model portion in the stage sequence using the image features, and to transmit gradients of the image features back to the video encoder in the stage sequence during the training process; A second training module is configured to train the video encoder in the stage sequence using the image data and the gradient of the image features.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause the computer to execute the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Fragmentation parallel model training method and device, equipment and storage medium
CN116704291A
Visual language model parameter alignment method and device, storage medium and electronic equipment
CN118379749A
Model pre-training method and device for multi-language task
CN119293514A
Signal lamp control model training method and device, electronic equipment and storage medium
CN119811107A
Data processing method and apparatus
WO2024041479A1