Training method, device, electronic device, computer program and storage medium

By component decomposing the multimodal model, clarifying resource requirements and dependencies, and distributing resource allocation, the problems of slow training speed and low resource utilization of multimodal model are solved, and a more efficient training process is achieved.

CN119292785BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411621000.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-05-27
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

During the training process of existing multimodal models, the training speed is slow and the computing resource utilization rate is low.

Method used

By determining the training resource requirements and component dependencies of each model component of the multimodal model, a resource allocation plan is formulated based on the idle computing device resources and training resource requirements, and distributed model training is carried out according to the component dependencies and resource allocation plan.

Benefits of technology

It improves training speed and resource utilization, ensures that each model component operates in the most appropriate environment, improves hardware utilization efficiency, and realizes more efficient multimodal model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119292785B_ABST
    Figure CN119292785B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a training method, apparatus, electronic device, computer program, and storage medium. Embodiments of the present application can determine multiple model components that make up a multimodal model; based on the structural parameters of each model component, determine the training resource requirements and component dependency relationships of each model component; according to the idle computing device resources and training resource requirements, determine the resource allocation scheme corresponding to each model component; and allocate computing device resources to each model component according to the component dependency relationships and resource allocation scheme for distributed model training to obtain a trained multimodal model. In the embodiments of the present application, by clarifying the training resource requirements of each model component, the resource allocation of computing devices can be optimized, achieving a more efficient resource utilization rate, ensuring that each component can be trained in the most suitable environment, improving the training speed, and thus enhancing the overall training effect. Therefore, this solution can improve the speed and efficiency of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a training method, device, electronic device, computer program and storage medium. Background Art

[0002] Multimodal tasks refer to processing and analyzing data in multiple different modalities, such as images, text, audio, etc. Multimodal tasks have made significant progress in the field of artificial intelligence, and the research scope includes image description generation, video content understanding, text-to-image generation or image-to-text retrieval, etc. In order to process data in different formats, the current mainstream multimodal models usually use their own encoding and decoding components to parse each data format.

[0003] In the process of training a multimodal model, multiple trainings are usually performed. Each training will fix the parameters of some components in the model and only update the parameters of other components in the model. This is repeated many times until all components in the multimodal model are trained. Although this method is effective, it also leads to slow training speed and low utilization of computing resources. Summary of the invention

[0004] The embodiments of the present application provide a training method, device, electronic device, computer program and storage medium, which can improve the training speed and resource utilization.

[0005] The present application provides a model distributed training method, including:

[0006] Identify the multiple model components that make up the multimodal model;

[0007] Based on the structural parameters of each model component, determine the training resource requirements and component dependencies of each model component;

[0008] Determine the resource allocation plan for each model component based on idle computing device resources and training resource requirements;

[0009] According to the component dependencies and resource allocation scheme, computing device resources are allocated to each model component for distributed model training to obtain a trained multimodal model.

[0010] The present application also provides a model distributed training device, including:

[0011] A model unit, used for obtaining a multimodal model, wherein the multimodal model includes a plurality of model components;

[0012] The requirement unit, which determines the resource requirements of each model component;

[0013] An allocation unit, configured to allocate a corresponding device to each model component according to resource requirements based on device parameters of each device in the computing cluster;

[0014] The training unit is used to use the device to perform model training on the model components corresponding to the device to obtain a trained multimodal model.

[0015] In some embodiments, according to idle computing device resources and training resource requirements, a resource allocation scheme corresponding to each model component is determined, including:

[0016] For each model component, determine the component load and storage requirements corresponding to the model component;

[0017] Determine the resource allocation plan for each model component based on component load, storage requirements, and idle computing device resources.

[0018] In some embodiments, according to the component dependency and resource allocation scheme, computing device resources are allocated to each model component for distributed model training to obtain a trained multimodal model, including:

[0019] Based on component dependencies and resource allocation schemes, determine the parallel training components and serial training components of the model training task;

[0020] Generate resource allocation strategies for parallel training components and serial training components to perform training tasks according to the resource allocation plan;

[0021] The parallel training components and serial training components are distributedly trained according to component dependencies and resource allocation strategies to obtain a trained multimodal model.

[0022] In some embodiments, the parallel training component includes at least two encoding components and at least two decoding components; the serial training component includes a main component; and the parallel training component and the serial training component are distributedly trained according to the component dependency and resource allocation strategy to obtain a trained multimodal model, including:

[0023] According to the resource configuration strategy, a coding training device and a subject training device are respectively configured for at least two coding components and a subject component;

[0024] The coding component is trained collaboratively by using the coding training device and the subject training device to obtain the trained coding component;

[0025] configuring a decoding training device for at least two decoding components according to a resource configuration strategy;

[0026] The encoding training device, the main body training device and the decoding training device are used to collaboratively train the main body component and the decoding component to obtain a trained main body component and a trained decoding component;

[0027] Based on the trained encoding component, the trained main component and the trained decoding component, a trained multimodal model is obtained.

[0028] In some embodiments, the coding component is trained collaboratively using a coding training device and a subject training device to obtain a trained coding component, including:

[0029] By using the collaboration of the coding training device and the main body training device, the parameter update of the main body component is frozen, the training data is input into at least two coding components in parallel, the output of at least two coding components is used as the input of the main body component, and the at least two coding components are trained to obtain the trained coding components.

[0030] In some embodiments, the encoding training device, the subject training device and the decoding training device are used to collaboratively train the subject component and the decoding component to obtain the trained subject component and the trained decoding component, including:

[0031] By using the coordination of the encoding training device, the main body training device and the decoding training device, the trained encoding component is frozen, the training data is input into at least two encoding components in parallel, the output of at least two encoding components is used as the input of the main body component, the output of the main body component is used as the parallel input of at least two decoding components, the main body component and the at least two decoding components are trained to obtain the trained main body component and the trained decoding component.

[0032] In some embodiments, the coding training device and the subject training device are used in coordination, the parameter update of the subject component is frozen, the training data is input into at least two coding components in parallel, the output of the at least two coding components is used as the input of the subject component, the at least two coding components are trained, and the trained coding components are obtained, including:

[0033] In the coding training device, training data is input in parallel into at least two coding components to obtain outputs of at least two coding components;

[0034] Freeze the parameter update of the main component, in the main training device, use the output of at least two encoding components as the input of the main component to obtain the output of the main component;

[0035] In the coding training device, the parameters of the coding component are iteratively updated according to the output of the main component until the iteration termination condition is met to obtain the trained coding component.

[0036] In some embodiments, the coding training device, the subject training device and the decoding training device are used in coordination, the trained coding component is frozen, the training data is input into at least two coding components in parallel, the output of the at least two coding components is used as the input of the subject component, the output of the subject component is used as the parallel input of at least two decoding components, the subject component and the at least two decoding components are trained, and the trained subject component and the trained decoding component are obtained, including:

[0037] Freeze the parameter update of the trained coding component, in the coding training device, input the training data into at least two coding components in parallel to obtain outputs of at least two coding components;

[0038] In the subject training device, the outputs of at least two encoding components are used as inputs of the subject component to obtain an output of the subject component;

[0039] In a decoding training device, the output of the main component is input in parallel into at least two decoding components to obtain outputs of at least two decoding components;

[0040] In the main body training device and the decoding training device respectively, the parameters of the main body component and at least two decoding components are iteratively updated according to the outputs of the two decoding components until the iteration termination condition is met, thereby obtaining a trained main body component and a trained decoding component.

[0041] In some embodiments, training a main component and at least two decoding components to obtain a trained main component and a trained decoding component includes:

[0042] Freeze the parameter update of the main component, introduce the main additional parameters, iteratively update the main additional parameters and the parameters of at least two decoding components until the iteration termination condition is met, and obtain the trained decoding component and the main additional parameters after the iteration termination;

[0043] The main extra parameters and main components after the iteration termination are used as the trained main components.

[0044] In some embodiments, before configuring a decoding training device for at least two decoding components according to a resource configuration strategy, the method further includes:

[0045] Unloading at least two decoding components from the multimodal model and storing structural parameters of at least two decoding groups on a hard disk of the device;

[0046] Before configuring a decoding training device for at least two decoding components according to a resource configuration strategy, the method further includes:

[0047] Read structural parameters of at least two decoding components from a hard disk of the device, and load the at least two decoding components into the multimodal model.

[0048] In some embodiments, configuring a decoding training device for at least two decoding components according to a resource configuration strategy includes:

[0049] Determine the resource update strategy corresponding to the trained encoding component;

[0050] Determine a shared training device from the encoding training devices based on a resource update strategy;

[0051] At least one of a shared training device and a decoding training device is configured for at least two decoding components according to a resource configuration strategy.

[0052] An embodiment of the present application also provides an electronic device, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in any one of the model distributed training methods provided in the embodiment of the present application.

[0053] An embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps in any one of the model distributed training methods provided in the embodiment of the present application.

[0054] The embodiments of the present application can determine multiple model components that constitute a multimodal model; based on the structural parameters of each model component, determine the training resource requirements and component dependencies of each model component; determine the resource allocation plan corresponding to each model component according to the idle computing device resources and training resource requirements; according to the component dependencies and the resource allocation plan, allocate computing device resources to each model component for distributed model training to obtain a trained multimodal model.

[0055] This application decouples each model component in the multimodal model, splits a complete model into multiple different components, optimizes the resource allocation of the computing cluster by clarifying the resource requirements of each model component, and thus re-matches different components with the equipment, solves the problem of uneven distribution of computing resources during the overall training process, which leads to slow training speed, and achieves more efficient training. The dynamic allocation based on device parameters ensures that each component can operate in the most suitable environment, improving the utilization efficiency of the hardware. In addition, assigning corresponding equipment to each component for model training can enable the training of all components to be executed in parallel, further accelerating the training process, thereby improving the overall training effect. The embodiments of this application are particularly effective in large-scale multimodal tasks, which can improve the training speed and resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1a It is a schematic diagram of the allocation relationship of the model distributed training method provided in the embodiment of the present application;

[0058] Figure 1b It is a flowchart of the model distributed training method provided in the embodiment of the present application;

[0059] Figure 1c It is a schematic diagram of the multimodal model structure of the model distributed training method provided in the embodiment of the present application;

[0060] Figure 1d This is a schematic diagram of the first stage of training of the model distributed training method provided in an embodiment of the present application;

[0061] Figure 1e This is a schematic diagram of the second stage training of the model distributed training method provided in an embodiment of the present application;

[0062] Figure 2 It is a structural diagram of a model distributed training device provided in an embodiment of the present application;

[0063] Figure 3 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0064] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0065] Embodiments of the present application provide a training method, apparatus, electronic device, computer program, and storage medium.

[0066] The model distributed training device can be integrated into an electronic device, which can be a terminal, a server, or other devices. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, or a personal computer (PC), etc. The server can be a single server or a server cluster composed of multiple servers.

[0067] In some embodiments, the model distributed training device can also be integrated into multiple electronic devices. For example, the model distributed training device can be integrated into multiple servers, and the model distributed training method of the present application can be implemented by multiple servers.

[0068] In some embodiments, the terminal may also be used as a server to implement part or all of the functions of the server.

[0069] For example, refer to Figure 1a , the electronic device can be a server, and the server can obtain a multimodal model, the multimodal model includes multiple model components, namely model component 1, model component 2...model component i+4; determine the resource requirement of each model component; based on the device parameters of each device in the computing cluster, allocate a corresponding device to each model component according to the resource requirement, for example, allocate model component i+1 and model component i+2 to device j+1; use the device to perform model training on the model component corresponding to the device to obtain a trained multimodal model.

[0070] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0071] In this embodiment, a model distributed training method is provided, such as Figure 1b As shown, the specific process of the model distributed training method can be as follows:

[0072] 110. Identify multiple model components that constitute a multimodal model.

[0073] Among them, multimodal models refer to neural network models that can process and understand multiple types (or modalities) of data, which can include text, images, audio, video, etc. Multimodal models can extract features from these different modal information and understand them to perform more complex tasks.

[0074] The modal model includes multiple model components (sub-models), each of which may be used to process data of a specific modality, so their computing resource consumption will be different during training. For example, it includes encoder components, decoder components, and backbone components.

[0075] Among them, the main component is the core part of the model, responsible for processing and fusing data from different modalities. For example, the main component can include a neural network with a Transformer architecture, a convolutional neural network (CNN), and so on. For example, the main component can be a model such as BERT, GPT, etc. For example, the main component can be an LLM (Large Language Model).

[0076] The encoding component can extract useful features from the input data, such as converting the input data into a feature vector; for example, the encoding component can include a video encoding component, an audio encoding component, an image encoding component, a text encoding component, etc. For example, the text encoding component can be a recurrent neural network (RNN), such as LSTM and GRU, etc.; for example, the image encoding component can be a convolutional neural network (CNN), etc.

[0077] The decoding component is used to generate the final output according to the feature vector, such as the translated text, the generated image description, etc. For example, the decoding component may include a video decoding component, an audio decoding component, an image decoding component, a text decoding component, etc.

[0078] For example, refer to Figure 1c The multimodal model includes an encoding component, a main component and a decoding component, wherein the encoding component includes an image encoding component, an audio encoding component and a video encoding component, the main component is LLM, and the decoding component includes an image decoding component, an audio decoding component and a video decoding component.

[0079] 120. Based on the structural parameters of each model component, determine the training resource requirements and component dependencies of each model component.

[0080] Among them, training resource requirements refer to the computing resources consumed by training the model component, such as component load, storage requirements, etc., and component load refers to the actual consumption of computing resources by the component during the training process. For example, the amount of computing (FLOPs), memory usage, and video memory usage required for a layer in forward propagation and backpropagation. The higher the load, the heavier the computing task undertaken by the component in the overall model.

[0081] Among them, component dependency refers to the input and output relationship between components. For example, assuming that the model consists of an encoding component, a main component, and a decoding component, the encoding component must complete the calculation first, and then the main component can use its output as input, and finally the decoding component can use the output of the main component to calculate the final task. The dependency relationship of these three components can be described as:

[0082] Encoding component: responsible for converting raw input (such as images or text) into feature representation. The encoding component is the starting point of the entire process, so there is no pre-dependency. The components that depend on the encoding component are the main components.

[0083] Main component: receives the feature representation generated by the encoding component and further processes these features (such as adding contextual information or enhancing the abstract representation of features). The main component relies on the output of the encoding component as input. The component that relies on the main component is the decoding component.

[0084] Decoding component: responsible for converting the abstract features generated by the main component into the final output (such as classification results or generated text, etc.). The encoding component is the end point of the entire process, so there is no post-dependency. The decoding component depends on the output of the main component.

[0085] Step 120 can determine the component load of each model component according to the model component structure parameters by reading, calculating or estimating. For example, the larger the number of model component parameters and the deeper the model layers, the more computing resources are required. Therefore, in some embodiments, the memory requirement consumed by training the model component can be estimated by the number of model component parameters.

[0086] For example, computing resources can be estimated through the model components' own properties, such as the number of model layers, convolution kernel size, number of channels, feature map size, complexity, etc.

[0087] For example, the computational effort of the convolutional layer in a model component can be calculated according to the formula:

[0088] Convolution FLOPs = 2 × (number of input channels × convolution kernel size × number of output channels × output feature map size)

[0089] For example, the computational effort of the fully connected layer in a model component can be calculated according to the formula:

[0090] Fully connected FLOPs = 2 × (number of input neurons × number of output neurons)

[0091] For example, the computational effort of a Transformer layer with an embedding dimension of d and a complexity of n2 in a model component can be calculated according to the formula:

[0092] Transformer FLOPs = 4 × n2 × d

[0093] Finally, the total computational effort of the model components is obtained by summing up the convolution FLOPs, fully connected FLOPs, and Transformer FLOPs:

[0094] Total computation = convolution FLOPs + fully connected FLOPs + Transformer FLOPs

[0095] In some embodiments, the resource requirements for training different model components are different. For example, the computing resource requirements of image processing components are often relatively high, especially the GPU memory consumption and computing power (FLOPs); the GPU memory requirements of text processing components may be slightly lower, but for long texts, the computing resource consumption is still significant; the computing resource consumption of audio processing components depends on the length and sampling rate of the audio. When processing long audio data, especially in tasks that require real-time processing, the computing and GPU memory resource requirements are also large.

[0096] , determine the resource allocation plan corresponding to each model component based on the idle computing device resources and training resource requirements.

[0097] Among them, computing device resources can be distributed computing systems, cloud computing platforms, etc., and devices can be independent computing nodes in a computing cluster, physical servers, or virtual machines. Each device has its own hardware resource device parameters, such as: CPU, GPU, memory, video memory, hard disk storage, network bandwidth, etc.

[0098] Resource allocation refers to how to effectively allocate available computing resources (such as CPU, GPU, memory, etc.) to different components when performing computing or training tasks to meet their needs and optimize overall performance. A good resource allocation plan can ensure that each task runs efficiently within the available resources, avoid resource waste, and improve training speed and model performance.

[0099] In some embodiments, these devices can be managed and scheduled through a cluster scheduling system such as Kubernetes, Slurm, Ray, etc., and can dynamically allocate computing resources to support different deep learning tasks.

[0100] Different model components have different computing and memory requirements during training. Therefore, a reasonable resource allocation strategy can improve model training efficiency.

[0101] For example, all model components are numbered 1 to n. According to the resource requirements of all model components, such as computing demand C={C1,C2,…,Cn}, storage demand Mf={Mf1,Mf2,…,Mfn} in the frozen state, and storage demand M={M1,M2,…,Mn} in the normal state, corresponding devices can be allocated to each model component based on the total computing resources Γ and total storage resources Θ of the device to maximize the utilization efficiency of device resources.

[0102] For example, define an array dp[i][j][k], which represents the maximum computing load that can be achieved using j units of computing resources and k units of storage resources when considering the first i components. The size of this array is (n+1)×(Γ+1)×(Θ+1)(n+1). Traverse all components and fill the dp table according to the state transition equation. Finally, the value in dp[n][Γ][Θ]dp[n] obtained by traversal represents the optimal allocation plan of model components under given resources.

[0103] This embodiment converts the allocation problem of model components and device resources into a knapsack problem and uses dynamic programming to efficiently find the best solution for allocating model components in the device cluster. This process not only improves the efficiency of resource utilization, but also helps to reasonably allocate computing and storage resources in complex model training.

[0104] Therefore, in some embodiments, the training resource requirements may include component load and storage requirements. According to the idle computing device resources and the training resource requirements, the resource allocation scheme corresponding to each model component is determined, including:

[0105] For each model component, determine the component load and storage requirements corresponding to the model component;

[0106] Determine the resource allocation scheme for each model component based on component load, storage requirements, and idle computing device resources. Figure 1a , model components 1 and 2 can be assigned to device 1, model components 3 and 4 can be assigned to device 2, model component i can be assigned to device j, model components i+1 and i+2 can be assigned to device j+1, and model components i+3 and i+4 can be assigned to device j+2.

[0107] 140. According to the component dependency and resource allocation plan, computing device resources are allocated to each model component for distributed model training to obtain a trained multimodal model.

[0108] Distributed model training refers to the process of performing model training in parallel on multiple computing devices. Distributed training can improve training speed and efficiency, especially when dealing with large-scale data or complex models.

[0109] For example, in some embodiments, distributed model training can be divided into parallel training and serial training, that is, model components can be divided into parallel training components and serial training components. Parallel training components refer to components that can be trained simultaneously. Usually, since there is no strong dependency between components, they can be processed in parallel on different computing devices, thereby speeding up the overall training speed; serial training components are components that need to be trained in sequence. Usually, due to component dependencies, such as the output of one component is the input of another component, serial training may cause the training time to be extended, but it ensures that the dependencies between components are satisfied.

[0110] Therefore, step 140 includes:

[0111] Based on component dependencies and resource allocation schemes, determine the parallel training components and serial training components of the model training task;

[0112] Generate resource allocation strategies for parallel training components and serial training components to perform training tasks according to the resource allocation plan;

[0113] The parallel training components and serial training components are distributedly trained according to the component dependencies and resource allocation strategies to obtain the trained multimodal model. The resource allocation strategy is a specific configuration generated based on the resource allocation scheme, which is used to guide how the parallel and serial training components use the resources of the computing device when performing training tasks, and may include memory allocation, computing power allocation, etc.

[0114] In some embodiments, according to the component dependencies, the encoding component and the decoding component can be trained independently without waiting for the output of other components. Therefore, the parallel training component may include at least two encoding components and at least two decoding components; the main component needs to be trained in a specific order to meet the dependencies, so the serial training component may include the main component.

[0115] In some embodiments, each device may correspond to at least one model component. If the device corresponds to multiple model components, the parallel execution of the training tasks of the multiple model components can be planned by a greedy strategy.

[0116] For example, the component and resource allocation scheme obtained in step 1 may be: Encoder 1 is allocated to device A, Encoder 2 is allocated to device B, Decoder 1 is allocated to device C, and the main model (LLM) is allocated to device D. According to the computing load and storage requirements, the parallel strategy can be determined, for example, device A processes Encoder 1 and Encoder 2 simultaneously, device B independently processes Decoder 1, and device C independently processes the main model (LLM).

[0117] Stage-wise training can gradually optimize different model components. Each stage can reduce computational complexity by fixing (freezing) the parameters of certain model components and focus on optimizing other components. Fixed parameters mean that the weights of these components are not updated during training, that is, the gradients of these weights are not calculated, thus saving computing resources and maintaining the stability of the model to prevent overfitting or over-adjustment of certain model components.

[0118] For example, the encoding component is trained in the first stage, and the main component and the decoding component are trained in the second stage.

[0119] Therefore, according to the component dependencies and resource allocation strategies, the parallel training components and serial training components are distributedly trained to obtain the trained multimodal model, including:

[0120] According to the resource configuration strategy, a coding training device and a subject training device are respectively configured for at least two coding components and a subject component;

[0121] The coding component is trained collaboratively by using the coding training device and the subject training device to obtain the trained coding component;

[0122] configuring a decoding training device for at least two decoding components according to a resource configuration strategy;

[0123] The encoding training device, the main body training device and the decoding training device are used to collaboratively train the main body component and the decoding component to obtain a trained main body component and a trained decoding component;

[0124] A trained multimodal model is obtained based on the trained encoding component, the trained main component and the trained decoding component.

[0125] Collaborative training refers to the process in which multiple devices work together to train certain components, which helps to make full use of resources and improve training efficiency.

[0126] Therefore, in some embodiments, the coding component is trained collaboratively using a coding training device and a subject training device to obtain a trained coding component, including:

[0127] By using the collaboration of the coding training device and the main body training device, the parameter update of the main body component is frozen, the training data is input into at least two coding components in parallel, the output of at least two coding components is used as the input of the main body component, and the at least two coding components are trained to obtain the trained coding components.

[0128] Among them, freezing the parameters of the main component can ensure that the main component remains stable when training the encoding component, avoiding changes in its parameters during the training process, helping to focus on optimizing the performance of the encoding component, reducing computational overhead, and preventing overfitting, ensuring that the training process is more efficient and effective.

[0129] Among them, training fixed (frozen) model components still consumes some computing resources, but the consumption is significantly less than that of model components in normal state. For example, training frozen model components does not involve back propagation and gradient updates, but still consumes some computing resources in the forward propagation stage.

[0130] In some embodiments, the encoding training device, the subject training device and the decoding training device are used to collaboratively train the subject component and the decoding component to obtain the trained subject component and the trained decoding component, including:

[0131] By using the coordination of the encoding training device, the main body training device and the decoding training device, the trained encoding component is frozen, the training data is input into at least two encoding components in parallel, the output of at least two encoding components is used as the input of the main body component, the output of the main body component is used as the parallel input of at least two decoding components, the main body component and the at least two decoding components are trained to obtain the trained main body component and the trained decoding component.

[0132] Among them, freezing the parameters of the trained encoding component can ensure that the trained encoding component remains stable when training the main component and the decoding component, avoiding changes in its parameters during the training process, helping to focus on optimizing the performance of the main component and the decoding component, reducing computational overhead, and preventing overfitting, ensuring that the training process is more efficient and effective.

[0133] In some embodiments, the coding training device and the subject training device are used in coordination, the parameter update of the subject component is frozen, the training data is input into at least two coding components in parallel, the output of the at least two coding components is used as the input of the subject component, the at least two coding components are trained, and the trained coding components are obtained, including:

[0134] In the coding training device, training data is input in parallel into at least two coding components to obtain outputs of at least two coding components;

[0135] Freeze the parameter update of the main component, in the main training device, use the output of at least two encoding components as the input of the main component to obtain the output of the main component;

[0136] In the coding training device, the parameters of the coding component are iteratively updated according to the output of the main component until the iteration termination condition is met to obtain the trained coding component.

[0137] Among them, because multiple encoding components can process data simultaneously, parallel input can significantly improve training efficiency.

[0138] In some embodiments, the coding training device, the subject training device and the decoding training device are used in coordination, the trained coding component is frozen, the training data is input into at least two coding components in parallel, the output of the at least two coding components is used as the input of the subject component, the output of the subject component is used as the parallel input of at least two decoding components, the subject component and the at least two decoding components are trained, and the trained subject component and the trained decoding component are obtained, including:

[0139] Freeze the parameter update of the trained coding component, in the coding training device, input the training data into at least two coding components in parallel to obtain outputs of at least two coding components;

[0140] In the subject training device, the outputs of at least two encoding components are used as inputs of the subject component to obtain an output of the subject component;

[0141] In a decoding training device, the output of the main component is input in parallel into at least two decoding components to obtain outputs of at least two decoding components;

[0142] In the main body training device and the decoding training device respectively, the parameters of the main body component and at least two decoding components are iteratively updated according to the outputs of the two decoding components until the iteration termination condition is met, thereby obtaining a trained main body component and a trained decoding component.

[0143] By inputting training data to multiple encoding components in parallel, the efficiency of feature extraction is improved and the training time is shortened. At the same time, freezing the parameters of the trained encoding components keeps the main component stable during the training process, avoiding interference caused by parameter updates and promoting convergence speed and stability. The output of the main component is used as the input of the decoding component, ensuring that the decoding process is based on the latest features, thereby enhancing the expressiveness of the model. In addition, the collaborative training mechanism iteratively updates the parameters of the main and decoding components to ensure that each component can dynamically adapt to each other's needs, thereby improving the performance of the overall model.

[0144] In some embodiments, training a main component and at least two decoding components to obtain a trained main component and a trained decoding component includes:

[0145] Freeze the parameter update of the main component, introduce the main additional parameters, iteratively update the main additional parameters and the parameters of at least two decoding components until the iteration termination condition is met, and obtain the trained decoding component and the main additional parameters after the iteration termination;

[0146] The main extra parameters and main components after the iteration termination are used as the trained main components.

[0147] Among them, the main additional parameters refer to the additional trainable parameters introduced in the main component, the purpose of which is to enhance the expressiveness or adaptability of the model. For example, the main additional parameters introduced can be additional fully connected layers, feature transformation matrices, low-rank matrices, regularization parameters, etc. added to the main component.

[0148] For example, the main components can be adaptively fine-tuned by introducing low-rank matrices. By introducing low-rank matrices, the number of parameters that need to be trained is significantly reduced, thereby reducing memory usage and computational burden.

[0149] In some embodiments, the configuration of device resources can be dynamically adjusted when switching between the two stages, that is, devices that can be shared are identified from the encoding training devices and transferred to the decoding training devices, ensuring optimal utilization of device resources and making switching between encoding and decoding tasks more efficient.

[0150] Therefore, in some embodiments, configuring a decoding training device for at least two decoding components according to a resource configuration strategy includes:

[0151] Determine the resource update strategy corresponding to the trained encoding component;

[0152] Determine a shared training device from the encoding training devices based on a resource update strategy;

[0153] At least one of a shared training device and a decoding training device is configured for at least two decoding components according to a resource configuration strategy.

[0154] In some embodiments, incorporating unloading and loading steps of the decoding component can more efficiently manage computing resources.

[0155] For example, in order to ensure that the decoding component does not occupy device resources, before configuring the decoding training device for at least two decoding components according to the resource configuration strategy, the method further includes:

[0156] Unload at least two decoding components from the multimodal model, store the structural parameters of at least two decoding groups on the device hard disk to free up computing and storage resources, and ensure that the encoding component is not disturbed during training.

[0157] Before configuring a decoding training device for at least two decoding components according to the resource configuration strategy, the method further includes:

[0158] Read structural parameters of at least two decoding components from a hard disk of the device, and load the at least two decoding components into the multimodal model to ensure that the decoding components can be quickly accessed and used.

[0159] For example, refer to Figure 1d Before the first stage of training, the speech decoding component, image decoding component, video decoding component, and text decoding component are unloaded from the multimodal model and stored in the memory; the first stage of training can use device 1 and device N+1 to perform the first stage training on the speech encoding component and image encoding component in the model component, and use device 2 and device N+2 to perform the first stage training on the video encoding component and alignment component in the model component, and use device N and device N-1 to perform the first stage training on the frozen main component, so as to obtain the speech, image, video encoding components and alignment components after the first stage of training.

[0160] Then, refer to Figure 1e Before the second stage of training, the speech decoding component, image decoding component, video decoding component, and text decoding component are read from the device's memory and loaded into the multimodal model; the second stage of training can determine device N+1 and device N+2 as shared devices, and use them for speech, image, video, and text decoding components. Ultimately, the second stage of training can use device 1 and device 2 to train the frozen speech, image, and video encoding components and alignment components, use device N+1 to train the speech and image decoding components, and use device N+2 to train the video and text decoding components.

[0161] In some embodiments, to further improve training efficiency, step 140 further includes:

[0162] Determine the load on each device in the computing cluster while training the model;

[0163] Based on the load of each device, reallocate the corresponding device to each model component according to resource requirements.

[0164] By evaluating the current load of each device in the computing cluster, the allocation of devices can be dynamically adjusted to effectively utilize its resources. This flexible resource management strategy can significantly improve the efficiency of model training, ensure that each component can learn in the optimal environment, and maximize the potential of computing resources.

[0165] As can be seen from the above, the embodiment of the present application can obtain a multimodal model, which includes multiple model components; determine the resource requirements of each model component; allocate corresponding devices to each model component according to resource requirements based on the device parameters of each device in the computing cluster; use the device to perform model training on the model component corresponding to the device to obtain a trained multimodal model. Therefore, this solution can effectively improve the efficiency and effectiveness of multimodal model training through systematic model component management, accurate resource demand assessment and flexible device allocation strategy. By achieving optimal resource utilization and load balancing, this method not only improves the training speed, but also provides a good foundation for the expansion and adaptation of the model, thereby improving the training speed and resource utilization.

[0166] The method described in the above embodiment will be further described in detail below.

[0167] In this embodiment, the method of the embodiment of the present application will be described in detail by taking the stage training of the multimodal model as an example.

[0168] 1. Decoupling of multimodal models.

[0169] The components of multiple different stages in the multimodal model are separated. After evaluating the computing and storage resource requirements of each component based on the model, the sequential dependencies between the components are decoupled, and the component tasks are reorganized and divided into corresponding computing resources for training.

[0170] The components of the multimodal model include the Encoder component, the Decoder component corresponding to different inputs, and the main LLM model component.

[0171] 2. Multi-stage training strategy.

[0172] (1) Component requirements assessment.

[0173] Evaluate the computational effort C of each component, the storage requirement Mf in the frozen state, and the storage requirement M in the normal state.

[0174] All components are numbered in order of dependency, and the computing requirements of all components are recorded as C={C1,C2,…,Cn}, the storage requirements in the frozen state are Mf={Mf1,Mf2,…,Mfn}, and the storage requirements in the normal state are M={M1,M2,…,Mn}.

[0175] (2) Task allocation.

[0176] Based on the total computing resources Γ and total storage resources Θ of the device, a corresponding device is assigned to each model component to maximize the utilization efficiency of the device resources.

[0177] The task allocation problem can be simplified as a knapsack problem, and dynamic programming is used to solve the optimal allocation of model components and equipment resources.

[0178] (3) Parallel strategy decision-making.

[0179] Training of multiple components can be scheduled to execute in parallel within each device.

[0180] For example, a greedy strategy is used to search in the existing parallel strategy space to obtain the parallel strategy configuration corresponding to the current task load.

[0181] (4) Model training.

[0182] Distributed training is performed based on task allocation and parallel strategy decisions.

[0183] 3. Tidal scheduling strategy for training tasks.

[0184] For example, multi-stage training may lead to resource imbalance, and the following strategies are proposed:

[0185] Phase 1: Train all Encoder components, freeze LLM parameters, and do not participate in the calculation of Decoder.

[0186] Phase 2: Freeze all encoders, use Lora to train LLM, and train all decoders.

[0187] Before the task starts, the training method of each stage is obtained, and the resource allocation strategy is determined based on the combined optimization strategy. When switching between each training stage, dynamic reallocation is performed according to the determined resource allocation strategy. Tidal scheduling improves resource utilization, reduces the IO overhead of model migration, and avoids manual intervention.

[0188] This solution aims to improve the training efficiency of the multimodal model through fine task allocation and resource scheduling, ensure resource balance between different stages, and thus achieve more efficient model training and resource utilization. As can be seen from the above, the embodiments of the present application can improve the training speed and resource utilization.

[0189] It can be understood that in the specific implementation of this application, relevant data such as equipment parameters and model training data are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0190] In order to better implement the above method, the embodiment of the present application also provides a model distributed training device, which can be integrated in an electronic device, and the electronic device can be a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers.

[0191] For example, in this embodiment, the method of the embodiment of the present application will be described in detail by taking the specific integration of the model distributed training device in the server as an example.

[0192] For example, Figure 2 As shown, the model distributed training device may include a model unit 210, a demand unit 220, an allocation unit 230, and a training unit 240, as follows:

[0193] (a) Model unit 210.

[0194] The model unit 210 is used to obtain a multimodal model, and the multimodal model includes multiple model components.

[0195] (ii) Demand unit 220.

[0196] The requirement unit 220 is used to determine the training resource requirements and component dependencies of each model component based on the structural parameters of each model component.

[0197] (iii) Allocation unit 230.

[0198] The allocation unit 230 is used to determine the resource allocation scheme corresponding to each model component according to the idle computing device resources and the training resource requirements.

[0199] (iv) Training unit 240.

[0200] The training unit 240 is used to allocate computing device resources to each model component for distributed model training according to component dependencies and resource allocation plans to obtain a trained multimodal model.

[0201] In some embodiments, according to idle computing device resources and training resource requirements, a resource allocation scheme corresponding to each model component is determined, including:

[0202] For each model component, determine the component load and storage requirements corresponding to the model component;

[0203] Determine the resource allocation plan for each model component based on component load, storage requirements, and idle computing device resources.

[0204] In some embodiments, according to the component dependency and resource allocation scheme, computing device resources are allocated to each model component for distributed model training to obtain a trained multimodal model, including:

[0205] Based on component dependencies and resource allocation schemes, determine the parallel training components and serial training components of the model training task;

[0206] Generate resource allocation strategies for parallel training components and serial training components to perform training tasks according to the resource allocation plan;

[0207] The parallel training components and serial training components are distributedly trained according to component dependencies and resource allocation strategies to obtain a trained multimodal model.

[0208] In some embodiments, the parallel training component includes at least two encoding components and at least two decoding components; the serial training component includes a main component; and the parallel training component and the serial training component are distributedly trained according to the component dependency and resource allocation strategy to obtain a trained multimodal model, including:

[0209] According to the resource configuration strategy, a coding training device and a subject training device are respectively configured for at least two coding components and a subject component;

[0210] The coding component is trained collaboratively by using the coding training device and the subject training device to obtain the trained coding component;

[0211] configuring a decoding training device for at least two decoding components according to a resource configuration strategy;

[0212] The encoding training device, the main body training device and the decoding training device are used to collaboratively train the main body component and the decoding component to obtain a trained main body component and a trained decoding component;

[0213] A trained multimodal model is obtained based on the trained encoding component, the trained main component and the trained decoding component.

[0214] In some embodiments, the coding component is trained collaboratively using a coding training device and a subject training device to obtain a trained coding component, including:

[0215] By using the collaboration of the coding training device and the main body training device, the parameter update of the main body component is frozen, the training data is input into at least two coding components in parallel, the output of at least two coding components is used as the input of the main body component, and the at least two coding components are trained to obtain the trained coding components.

[0216] In some embodiments, the encoding training device, the subject training device and the decoding training device are used to collaboratively train the subject component and the decoding component to obtain the trained subject component and the trained decoding component, including:

[0217] By using the coordination of the encoding training device, the main body training device and the decoding training device, the trained encoding component is frozen, the training data is input into at least two encoding components in parallel, the output of at least two encoding components is used as the input of the main body component, the output of the main body component is used as the parallel input of at least two decoding components, the main body component and the at least two decoding components are trained to obtain the trained main body component and the trained decoding component.

[0218] In some embodiments, the coding training device and the subject training device are used in coordination, the parameter update of the subject component is frozen, the training data is input into at least two coding components in parallel, the output of the at least two coding components is used as the input of the subject component, the at least two coding components are trained, and the trained coding components are obtained, including:

[0219] In the coding training device, training data is input in parallel into at least two coding components to obtain outputs of at least two coding components;

[0220] Freeze the parameter update of the main component, in the main training device, use the output of at least two encoding components as the input of the main component to obtain the output of the main component;

[0221] In the coding training device, the parameters of the coding component are iteratively updated according to the output of the main component until the iteration termination condition is met to obtain the trained coding component.

[0222] In some embodiments, the coding training device, the subject training device and the decoding training device are used in coordination, the trained coding component is frozen, the training data is input into at least two coding components in parallel, the output of the at least two coding components is used as the input of the subject component, the output of the subject component is used as the parallel input of at least two decoding components, the subject component and the at least two decoding components are trained, and the trained subject component and the trained decoding component are obtained, including:

[0223] Freeze the parameter update of the trained coding component, in the coding training device, input the training data into at least two coding components in parallel to obtain outputs of at least two coding components;

[0224] In the subject training device, the outputs of at least two encoding components are used as inputs of the subject component to obtain an output of the subject component;

[0225] In a decoding training device, the output of the main component is input in parallel into at least two decoding components to obtain outputs of at least two decoding components;

[0226] In the main body training device and the decoding training device respectively, the parameters of the main body component and at least two decoding components are iteratively updated according to the outputs of the two decoding components until the iteration termination condition is met, thereby obtaining a trained main body component and a trained decoding component.

[0227] In some embodiments, training a main component and at least two decoding components to obtain a trained main component and a trained decoding component includes:

[0228] Freeze the parameter update of the main component, introduce the main additional parameters, iteratively update the main additional parameters and the parameters of at least two decoding components until the iteration termination condition is met, and obtain the trained decoding component and the main additional parameters after the iteration termination;

[0229] The main extra parameters and main components after the iteration termination are used as the trained main components.

[0230] In some embodiments, before configuring a decoding training device for at least two decoding components according to a resource configuration strategy, the method further includes:

[0231] Unloading at least two decoding components from the multimodal model and storing structural parameters of at least two decoding groups on a hard disk of the device;

[0232] Before configuring a decoding training device for at least two decoding components according to a resource configuration strategy, the method further includes:

[0233] Read structural parameters of at least two decoding components from a hard disk of the device, and load the at least two decoding components into the multimodal model.

[0234] In some embodiments, configuring a decoding training device for at least two decoding components according to a resource configuration strategy includes:

[0235] Determine the resource update strategy corresponding to the trained encoding component;

[0236] Determine a shared training device from the encoding training devices based on a resource update strategy;

[0237] At least one of a shared training device and a decoding training device is configured for at least two decoding components according to a resource configuration strategy.

[0238] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can refer to the previous method embodiments, which will not be repeated here.

[0239] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0240] As can be seen from the above, the model distributed training device of this embodiment determines the multiple model components constituting the multimodal model by the model unit; the demand unit determines the training resource requirements and component dependencies of each model component based on the structural parameters of each model component; the allocation unit determines the resource allocation scheme corresponding to each model component according to the idle computing device resources and the training resource requirements; the training unit allocates computing device resources to each model component according to the component dependencies and the resource allocation scheme for distributed model training to obtain the trained multimodal model. Therefore, the embodiment of the present application can improve the training speed and resource utilization.

[0241] The embodiment of the present application also provides an electronic device, which can be a terminal, a server, etc. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, a personal computer, etc. The server can be a single server or a server cluster composed of multiple servers, etc.

[0242] In some embodiments, the model distributed training device can also be integrated into multiple electronic devices. For example, the model distributed training device can be integrated into multiple servers, and the model distributed training method of the present application can be implemented by multiple servers.

[0243] In this embodiment, the electronic device of this embodiment is a server as an example for detailed description, for example, Figure 3 As shown, it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:

[0244] The electronic device may include components such as a processor 310 with one or more processing cores, a memory 320 with one or more computer-readable storage media, a power supply 330, an input module 340, and a communication module 350. Those skilled in the art will appreciate that Figure 3 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0245] The processor 310 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 320, and calling data stored in the memory 320, it executes various functions of the electronic device and processes data, thereby performing overall detection of the electronic device. In some embodiments, the processor 310 may include one or more processing cores; in some embodiments, the processor 310 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 310.

[0246] The memory 320 can be used to store software programs and modules. The processor 310 executes various functional applications and data processing by running the software programs and modules stored in the memory 320. The memory 320 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 320 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one memory storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 320 may also include a memory controller to provide the processor 310 with access to the memory 320.

[0247] The electronic device also includes a power supply 330 for supplying power to various components. In some embodiments, the power supply 330 can be logically connected to the processor 310 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 330 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.

[0248] The electronic device may further include an input module 340, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0249] The electronic device may further include a communication module 350. In some embodiments, the communication module 350 may include a wireless module. The electronic device may perform short-range wireless transmission through the wireless module of the communication module 350, thereby providing the user with wireless broadband Internet access. For example, the communication module 350 may be used to help the user send and receive emails, browse web pages, and access streaming media.

[0250] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 310 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 320 according to the following instructions, and the processor 310 will run the application programs stored in the memory 320, thereby realizing various functions, as follows:

[0251] Identify the multiple model components that make up the multimodal model;

[0252] Based on the structural parameters of each model component, determine the training resource requirements and component dependencies of each model component;

[0253] Determine the resource allocation plan for each model component based on idle computing device resources and training resource requirements;

[0254] According to the component dependencies and resource allocation scheme, computing device resources are allocated to each model component for distributed model training to obtain a trained multimodal model.

[0255] It can be seen from the above that the embodiments of the present application can improve the training speed and resource utilization.

[0256] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0257] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions, which can be loaded by a processor to execute the steps in any model distributed training method provided in the embodiment of the present application. For example, the instruction can execute the following steps:

[0258] Identify the multiple model components that make up the multimodal model;

[0259] Based on the structural parameters of each model component, determine the training resource requirements and component dependencies of each model component;

[0260] Determine the resource allocation plan for each model component based on idle computing device resources and training resource requirements;

[0261] According to the component dependencies and resource allocation scheme, computing device resources are allocated to each model component for distributed model training to obtain a trained multimodal model.

[0262] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a memory or an optical disk, etc.

[0263] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the methods provided in various optional implementations of the model training aspect or the multimodal task aspect provided in the above embodiments.

[0264] Since the instructions stored in the storage medium can execute the steps in any model distributed training method provided in the embodiments of the present application, the beneficial effects that can be achieved by any model distributed training method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0265] The above is a detailed introduction to a model distributed training method, device, electronic device and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. For technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A model distributed training method, characterized in that: include: Identify the multiple model components that make up the multimodal model; Based on the structural parameters of each model component, determine the training resource requirements and component dependencies of each model component; Determine a resource allocation scheme corresponding to each model component according to idle computing device resources and the training resource requirements; Based on the component dependency and the resource allocation scheme, determine a parallel training component and a serial training component of the model training task, wherein the parallel training component includes at least two encoding components and at least two decoding components, and the serial training component includes a main component; According to the resource allocation scheme, generating a resource configuration strategy for the parallel training component and the serial training component when performing training tasks; According to the resource configuration strategy, respectively configure a coding training device and a main body training device for the at least two coding components and the main body component; The coding training device and the subject training device are used to collaboratively train the coding component to obtain the trained coding component; configuring a decoding training device for the at least two decoding components according to the resource configuration strategy; The encoding training device, the subject training device and the decoding training device are used to collaboratively train the subject component and the decoding component to obtain a trained subject component and a trained decoding component; Based on the trained encoding component, the trained main component and the trained decoding component, a trained multimodal model is obtained.

2. The model distributed training method according to claim 1, characterized in that: The determining of a resource allocation scheme corresponding to each model component according to idle computing device resources and the training resource requirements includes: For each model component, determining a component load and storage requirement corresponding to the model component; Based on the component load, the storage requirement and the idle computing device resources, a resource allocation scheme corresponding to each model component is determined.

3. The model distributed training method according to claim 1, characterized in that: The method of using the coding training device and the subject training device to collaboratively train the coding component to obtain the trained coding component includes: The coding training device and the main body training device are used in collaboration to freeze the parameter update of the main body component, input the training data into at least two coding components in parallel, use the outputs of the at least two coding components as the input of the main body component, train the at least two coding components, and obtain the trained coding components.

4. The model distributed training method according to claim 1, characterized in that: The method of using the encoding training device, the subject training device and the decoding training device to collaboratively train the subject component and the decoding component to obtain the trained subject component and the trained decoding component includes: The encoding training device, the main body training device and the decoding training device are used in collaboration to freeze the trained encoding component, input the training data in parallel to at least two encoding components, use the outputs of the at least two encoding components as the input of the main body component, use the output of the main body component as the parallel input to the at least two decoding components, train the main body component and the at least two decoding components, and obtain a trained main body component and a trained decoding component.

5. The model distributed training method according to claim 1, characterized in that: The method adopts the coordination of the coding training device and the subject training device, freezes the parameter update of the subject component, inputs the training data into at least two coding components in parallel, uses the outputs of the at least two coding components as the inputs of the subject component, trains the at least two coding components, and obtains the trained coding components, including: In the coding training device, training data is input into at least two coding components in parallel to obtain outputs of the at least two coding components; Freeze the parameter update of the main component, and in the main training device, use the outputs of the at least two encoding components as inputs of the main component to obtain the output of the main component; In the coding training device, the parameters of the coding component are iteratively updated according to the output of the main component until an iteration termination condition is met to obtain a trained coding component.

6. The model distributed training method according to claim 3, characterized in that: The method adopts the coordination of the encoding training device, the subject training device and the decoding training device, freezes the trained encoding component, inputs the training data in parallel to at least two encoding components, uses the outputs of the at least two encoding components as the inputs of the subject component, uses the outputs of the subject component as the parallel inputs to the at least two decoding components, trains the subject component and the at least two decoding components, and obtains the trained subject component and the trained decoding component, including: Freezing parameter updates of the trained encoding components, in the encoding training device, inputting training data in parallel into at least two encoding components to obtain outputs of the at least two encoding components; In the subject training device, the outputs of the at least two encoding components are used as inputs of the subject component to obtain the output of the subject component; In the decoding training device, the output of the main component is input into at least two decoding components in parallel to obtain the outputs of the at least two decoding components; In the main body training device and the decoding training device respectively, the parameters of the main body component and the at least two decoding components are iteratively updated according to the outputs of the two decoding components until an iteration termination condition is met, thereby obtaining a trained main body component and a trained decoding component.

7. The model distributed training method according to claim 3, characterized in that: The step of training the main component and the at least two decoding components to obtain a trained main component and a trained decoding component includes: Freeze the parameter update of the main component, introduce the main additional parameters, iteratively update the main additional parameters and the parameters of the at least two decoding components until the iteration termination condition is met, and obtain the trained decoding component and the main additional parameters after the iteration termination, wherein the main additional parameters refer to the additional trainable parameters introduced in the main component; The additional parameters of the subject after the iteration is terminated and the subject component are used as the trained subject component.

8. The model distributed training method according to claim 1, characterized in that: Before configuring the decoding training device for the at least two decoding components according to the resource configuration strategy, the method further includes: Unloading the at least two decoding components from the multimodal model and storing the structural parameters of the at least two decoding groups on a device hard disk; Before configuring the decoding training device for the at least two decoding components according to the resource configuration strategy, the method further includes: The structural parameters of the at least two decoding components are read from the hard disk of the device, and the at least two decoding components are loaded into the multimodal model.

9. The model distributed training method according to claim 1, characterized in that: The configuring a decoding training device for the at least two decoding components according to the resource configuration strategy includes: Determine the resource update strategy corresponding to the trained encoding component; Determining a shared training device from the coding training devices based on the resource update strategy; At least one of the shared training device and the decoding training device is configured for the at least two decoding components according to the resource configuration strategy.

10. A model distributed training device, characterized in that: include: A model unit, which is used to determine a plurality of model components constituting a multimodal model; A requirement unit, which is used to determine the training resource requirements and component dependencies of each model component based on the structural parameters of each model component; An allocation unit, configured to determine a resource allocation scheme corresponding to each model component according to idle computing device resources and the training resource requirements; A training unit, configured to determine a parallel training component and a serial training component of a model training task based on the component dependency and the resource allocation scheme, wherein the parallel training component includes at least two encoding components and at least two decoding components, and the serial training component includes a main component; According to the resource allocation scheme, generating a resource configuration strategy for the parallel training component and the serial training component when performing training tasks; According to the resource configuration strategy, respectively configure a coding training device and a main body training device for the at least two coding components and the main body component; The coding training device and the subject training device are used to collaboratively train the coding component to obtain the trained coding component; configuring a decoding training device for the at least two decoding components according to the resource configuration strategy; The encoding training device, the subject training device and the decoding training device are used to collaboratively train the subject component and the decoding component to obtain a trained subject component and a trained decoding component; Based on the trained encoding component, the trained main component and the trained decoding component, a trained multimodal model is obtained.

11. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the steps in the model distributed training method described in any one of claims 1 to 9.

12. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores multiple instructions; the processor loads instructions from the memory to execute the steps in the model distributed training method described in any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores multiple instructions, which are suitable for loading by a processor to execute the steps in the model distributed training method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Resource management method and device, equipment and storage medium

    CN113467922A

  • Distributed training method, device and equipment based on end-to-end self-adaption

    CN114169427A

  • Distributed transformer substation defect detection-based multi-modal algorithm training method and device

    CN117972490A

  • Large model training method and device, equipment and storage medium

    CN118428486A