Model training optimization method and device, electronic equipment, storage medium and program
By acquiring information on model training and inference time consumption, the number of devices can be dynamically adjusted to achieve asynchronous pipelined training with separate training and inference, thus solving the problem of low training efficiency of neural network models and improving hardware resource utilization and training efficiency.
Patent Information
- Application Number
- CN202511099939.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-07
AI Technical Summary
In existing neural network models, the inference-coupled training method results in low training efficiency, low utilization of hardware resources, and waste of computing resources.
By obtaining the model training time and model inference time of the target optimization model, the target ratio of training instances to inference instances is determined, and the number of devices is dynamically adjusted to achieve asynchronous pipelined training with separate training and inference, thereby optimizing hardware resource configuration.
This improved the utilization of hardware resources in the coupled training process of the model, reduced latency, and improved overall training efficiency.
Smart Images

Figure CN120598063B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a model training optimization method, apparatus, electronic device, storage medium and program. Background Technology
[0002] In the training process of some types of neural network models, the model's inference and training are tightly coupled. That is, in each training step of the model, forward inference computation is required to obtain the model's output. Then, the input and output of this inference are combined as the training input, which is fed into the model for forward computation and backpropagation algorithm to update the model's parameters.
[0003] Currently, there are two main approaches to implementing inference-coupled training (or inference-training coupling) in neural network models. One approach is to directly perform inference using training instances, and then update the model by calculating the reward and output probability based on the generated content. The other approach is to use separate instances for inference, synchronizing with the training model before inference and distributing the results after inference.
[0004] In the process of developing this invention, the inventors discovered the following shortcomings in the existing technology: In the first inference-training coupling training method, because the inference speed is slower than the training speed, the training process is forced to wait, thus prolonging the overall training time. It also involves frequent training-inference switching, and the inference during training is not specifically optimized, resulting in low overall model training efficiency. In the second inference-training coupling training method, the training card is idle while waiting for inference data, and the inference card is also idle while waiting for training. Therefore, a synchronous scheduling method is adopted, causing the training card and inference card to alternately wait, resulting in a significant waste of computing resources. Furthermore, some model training methods focus on strategies such as data preprocessing and collection, improving model training performance, and optimization for specific domain applications. However, these methods mainly involve the application of models or large models in vertical domains or traditional model training methods, without addressing the inference-integrated inference-training coupling training process. They lack an overall consideration of the training framework, especially failing to solve the efficiency problem in the integrated inference-training coupling training process. Summary of the Invention
[0005] This invention provides a model training optimization method, apparatus, electronic device, storage medium, and program, which can improve the hardware resource utilization of the model during the push-training coupled training process, reduce the latency of push-training coupled training, and improve the efficiency of push-training coupled training.
[0006] According to one aspect of the present invention, a model training optimization method is provided, comprising:
[0007] The model time information of the target optimization model is obtained by using the device-associated hardware information of the target device running the target optimization model; wherein, the model time information includes the model training time and model inference time of the target optimization model;
[0008] The target ratio of training instances to inference instances of the target optimization model is determined based on the model training time and the model inference time.
[0009] The number of first target devices required for inference and the number of second target devices required for training of the target optimization model are determined based on the target ratio of training instances to inference instances of the target optimization model.
[0010] The target devices of the first target number and the target devices of the second target number are used to complete the asynchronous pipelined training process of the push-train separation of the target optimization model.
[0011] According to another aspect of the present invention, a model training optimization apparatus is provided, comprising:
[0012] The model time consumption information acquisition module is used to acquire the model time consumption information of the target optimization model by means of the device-associated hardware information of the target device running the target optimization model; wherein, the model time consumption information includes the model training time and model inference time of the target optimization model;
[0013] The inference-training instance ratio determination module is used to determine the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time.
[0014] The target device quantity determination module is used to determine the first target device quantity required for inference and the second target device quantity required for training of the target optimization model based on the target ratio of training instances to inference instances of the target optimization model.
[0015] The target devices of the first target number and the target devices of the second target number are used to complete the asynchronous pipelined training process of the push-train separation of the target optimization model.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the model training optimization method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the model training optimization method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the model training optimization method described in any embodiment of the present invention.
[0022] This invention obtains model time information, such as model training time and model inference time, from the device-associated hardware information of the target device running the target optimization model. Then, based on the model training time and model inference time, it determines the target ratio of training instances to inference instances of the target optimization model. Based on this ratio, it determines the number of first target devices required for inference and the number of second target devices required for training. After determining these numbers, the asynchronous pipelined training process of the target optimization model can be completed using the first and second target devices. This technical solution optimizes the resource allocation of the target device during model inference and training by using the device-associated hardware information of the target device. This significantly reduces the waiting time for computing card resources on the target device, solving the problems of low hardware resource utilization and low overall training efficiency in existing models using a push-train coupling training method. It improves the hardware resource utilization of the model during push-train coupling training, reduces the latency of push-train coupling training, and improves the efficiency of push-train coupling training.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a model training optimization method provided in an embodiment of the present invention;
[0026] Figure 2 This is a flowchart of another model training optimization method provided in an embodiment of the present invention;
[0027] Figure 3 This is a schematic diagram of a data coarse sorting process provided in an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram of a data sorting process provided in an embodiment of the present invention;
[0029] Figure 5 This is a schematic diagram of a process for dynamically determining the target ratio of training instances and inference instances according to an embodiment of the present invention;
[0030] Figure 6 This is a schematic diagram of the overall training process of performing asynchronous pipelined inference and training separation on a target optimization model according to an embodiment of the present invention;
[0031] Figure 7 This is a schematic diagram showing the comparison of the time consumption of different push-training coupling training methods provided in an embodiment of the present invention;
[0032] Figure 8 This is a schematic diagram of a model training optimization device provided in an embodiment of the present invention;
[0033] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] Large Language Models (LLMs) are deep learning models trained on massive amounts of relevant data (such as text, speech, or combined text and image data) to process text sequences. They can generate natural language text or understand the meaning of spoken text. These models typically have billions of parameters. LLMs can handle various natural language tasks, such as text classification, question answering, and dialogue, and have wide applications. The input to a LLM is data, such as text, speech, or combined text and image data. The LLM encodes the input data to obtain corresponding word vectors, and then decodes these word vectors to automatically process the input data and obtain the corresponding output data. For example, text can be input into a LLM, which then processes and predicts the input text, outputting the corresponding response text.
[0037] Taking large models as an example, the training of some types of large models includes pre-training and post-training processes. Post-training of large models typically uses reinforcement learning methods to improve model performance. Common reinforcement learning methods include Proximal Policy Optimization (PPO), Direct Preference Optimization (DPOD), and Group Reinforcement Policy Optimization (GRPO). In the training process of these large models, the model being trained is called the policy model. Some methods also require a given pre-trained model or a manually constructed model as a reference model. The training process for such large models first involves the policy model performing complete inference, generating responses for each input sample. Then, rewards and token probability distributions are calculated based on these responses, followed by the calculation of the loss function for training. The current coupled inference and training approach leads to significant performance overhead, severely impacting the overall training efficiency of the model.
[0038] Figure 1 This is a flowchart of a model training optimization method provided by an embodiment of the present invention. This embodiment is applicable to situations where model time consumption information is determined based on the device-associated hardware information of the model running device, and then the number of devices required for training and inference in the asynchronous pipelined training process of model training and inference is further determined based on the model time consumption information. This method can be executed by a model training optimization device, which can be implemented by software and / or hardware, and is generally integrated into an electronic device. The electronic device can be a terminal device or a server device, as long as it can execute the model training optimization method. The embodiments of the present invention do not limit the specific device type of the electronic device. Accordingly, as... Figure 1 As shown, the method includes the following operations:
[0039] S110. Obtain the model time information of the target optimization model by means of the device-associated hardware information of the target device running the target optimization model; wherein, the model time information includes the model training time and model inference time of the target optimization model.
[0040] The target optimization model can be a model type that uses an asynchronous pipelined training method with separate training and inference processes. For example, the target optimization model can be a reinforcement learning model, etc. This embodiment of the invention does not limit the specific model type of the target optimization model. Optionally, the target optimization model can be applied to various specific application scenarios. For example, when the target optimization model is a target tracking model, it can be applied to target tracking application scenarios. When the target optimization model is a speech recognition model or a text recognition model, it can be applied to speech recognition or text recognition application scenarios. The target device can be a device used for training and / or inference of the target optimization model, such as a server device with a computing card, etc. Device-associated hardware information can be hardware information associated with the target device, as long as it can be used to determine the reference model's time consumption information. This embodiment of the invention does not limit the hardware type or specific information content involved in the device-associated hardware information. Model training time can be the time consumed by the target optimization model in a single training process. Model inference time can be the time consumed by the target optimization model in a single inference process.
[0041] In this embodiment of the invention, before training the target optimization model using an asynchronous pipelined training method with separate inference and propagation, relevant hardware information related to the model training and inference processes in the target device used to run the target optimization model can be obtained first, serving as the device-associated hardware information of the target device. For example, the computing power and bandwidth of the inference card and training card of the target device can be obtained, serving as the device-associated hardware information of the target device. The inference card can be a computing card used by the target device for the inference process of the target optimization model, and the training card can be a computing card used by the target device for the training process of the target optimization model. A computing card, also known as a computing accelerator card, is a type of hardware specifically designed to perform high-performance computing tasks. It typically integrates a large number of computing units and high-speed memory to provide powerful parallel computing capabilities.
[0042] After obtaining the device-associated hardware information of the target device running the target optimization model, the model training time and model inference time of the target optimization model can be obtained based on this information. The model training time reflects the time distribution characteristics of the target optimization model's training, and the model inference time reflects the time distribution characteristics of the target optimization model's inference.
[0043] S120. Determine the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time.
[0044] Wherein, training instances can be target device instances used to train the target optimization model. Inference instances can be target device instances used to infer the target optimization model. The target ratio can be the ratio between the number of training instances and the number of inference instances. Optionally, the target ratio can be the optimal ratio between the number of inference instances and training instances.
[0045] After obtaining the training time and inference time of the target optimization model, the time distribution characteristics of the training and inference of the target optimization model are determined. Therefore, the target ratio of training instances to inference instances of the target optimization model can be further calculated based on the training time and inference time.
[0046] It should be noted that the target ratio of training instances to inference instances can be a fixed ratio or a dynamically updated ratio, and the embodiments of the present invention do not impose any restrictions on this.
[0047] S130. Determine the number of first target devices required for inference and the number of second target devices required for training of the target optimization model based on the target ratio of training instances to inference instances of the target optimization model; wherein, the number of target devices in the first target number and the number of target devices in the second target number are used to complete the asynchronous pipelined training process of the target optimization model with separate inference and training.
[0048] The first target device number can be the number of target devices used for inference of the target optimization model, and the second target device number can be the number of target devices used for training the target optimization model.
[0049] Understandably, the total number of target devices typically used to train the objective optimization model is known. Therefore, after determining the target ratio of training instances to inference instances in the overall training process, the number of first target devices required for inference and the number of second target devices required for training can be determined based on the total number of target devices and the calculated target ratio. The sum of the first and second target device numbers is the total number of target devices.
[0050] Optionally, the number of first and second target devices determined by the target ratio represents the optimal ratio between the inference cluster and the training cluster. This enables efficient allocation of hardware resources required for model training, as well as dynamic adjustment and utilization of device cluster resources during inference and training. By using the target devices of the first and second target device numbers to complete the asynchronous pipelined training process of the target optimization model, the waiting time of the target devices can be significantly reduced, thereby significantly reducing the overall training time of the target optimization model and improving the overall training efficiency of the model.
[0051] In an optional embodiment of the present invention, the asynchronous pipelined training process of the target optimization model is an interleaved execution process of the inference process and the training process; the first process in the asynchronous pipelined training process of the target optimization model is the inference process, and there is no inference waiting time between each inference process.
[0052] Accordingly, in the asynchronous coupled training process of the target optimization model, the inference process of the target optimization model can be executed by a first number of target devices, and the training process can be executed by a second number of target devices. The inference and training processes of the target optimization model are executed alternately, and the first step in the asynchronous coupled training process of the target optimization model is the inference process. That is, the asynchronous coupled training process of the target optimization model is specifically as follows: after the first number of target devices completes the inference process of the current round, the next round of inference process is executed immediately; and after the second number of target devices determines that the first number of target devices has completed the inference process of the current round, the current round of training process is executed. For example, after the first number of target devices completes the first round of inference, the second round of inference is immediately executed. The second number of target devices, after the first number of target devices completes the first round of inference, executes the first round of training. After the first number of target devices completes the second round of inference, the third round of inference is immediately executed. The second number of target devices, after the first number of target devices completes the second round of inference, executes the second round of training... and so on, until the overall asynchronous pipelined training process of the target optimization model is completed. Therefore, the asynchronous pipelined training process of the target optimization model is also a type of coupled training method.
[0053] This invention obtains model time information, such as model training time and model inference time, from the device-associated hardware information of the target device running the target optimization model. Then, based on the model training time and model inference time, it determines the target ratio of training instances to inference instances of the target optimization model. Based on this ratio, it determines the number of first target devices required for inference and the number of second target devices required for training. After determining these numbers, the asynchronous pipelined training process of the target optimization model can be completed using the first and second target devices. This technical solution optimizes the resource allocation of the target device during model inference and training by using the device-associated hardware information of the target device. This significantly reduces the waiting time for computing card resources on the target device, solving the problems of low hardware resource utilization and low overall training efficiency in existing models using a push-train coupling training method. It improves the hardware resource utilization of the model during push-train coupling training, reduces the latency of push-train coupling training, and improves the efficiency of push-train coupling training.
[0054] Figure 2 This is a flowchart of another model training optimization method provided by an embodiment of the present invention. This embodiment is based on the above embodiment and is further specified. In this embodiment, various specific optional implementation methods are given for obtaining the model time information of the target optimization model and determining the target ratio of training instances and inference instances of the target optimization model by associating hardware information of the device running the target optimization model. Correspondingly, as Figure 2 As shown, the method in this embodiment may include:
[0055] S210. In each training round of the target optimization model, data equalization preprocessing is performed on the training data of the target optimization model to obtain data preprocessing information.
[0056] Understandably, because the objective optimization model employs an asynchronous pipelined approach to execute the overall training process, each training epoch of the objective optimization model includes one inference process and one training process. The inference process uses the original training data as input data for model inference, while the training process concatenates the output of the corresponding epoch's inference process with the current original training data to obtain new training data, which is then used as input data for model training.
[0057] Optionally, the training data used to train the target optimization model consists of dialogue data, namely, input prompts and the model's responses. During the training process of the target optimization model, only the input data needs to be given to the model. First, the policy model corresponding to the target optimization model infers and generates a complete response based on the input prompt. Then, this input is concatenated with the corresponding output and fed into the policy model and reference model respectively for forward computation to obtain the original values or confidence scores (logits) of the model output for training. Therefore, the length of the input prompt directly affects the inference time and the forward computation time of subsequent training. Typically, prompts of similar length will produce outputs of similar length, but due to the randomness of the generated content, their length is not fixed. Within the same mini-batch, it is also necessary to ensure that each prompt is as consistent in length as possible. This reduces the number of tokens (the smallest unit of text processed by the model, which can be understood as a word, character, or symbol fragment) required for data storage, thereby reducing GPU memory consumption.
[0058] Therefore, in each inference and training round of the objective optimization model, the training data can first undergo data balancing preprocessing to ensure that the input data between each inference instance is as balanced as possible, thus making the inference time of each inference instance roughly the same; similarly, the input data between each training instance should be balanced as much as possible, thus making the training time of each training instance roughly the same. The data preprocessing information obtained from the data balancing preprocessing of the objective optimization model's training data can be used for subsequent time calculations.
[0059] Optionally, the data balancing preprocessing of the training data of the target optimization model can follow two preprocessing objectives. The first preprocessing objective is to ensure that the inference time of each inference instance is roughly the same, and to ensure that the total number of tokens of different prompts input to each inference instance is consistent. The second preprocessing objective is to ensure that the training time of each training instance is roughly the same, and to ensure that the token length of each mini-batch prompt of the training instance is consistent with the token length of the concatenated text generated by the target optimization model through the prompt.
[0060] In an optional embodiment of the present invention, the data balancing preprocessing of the training data of the target optimization model may include: determining the average number of tokens per inference instance based on the original training data of the target optimization model in the current training round; determining the original training sample data loaded by the data loader of the inference instance in the computing card based on the average number of tokens per inference instance, so as to achieve clustering processing of the inference sample data of the inference instance; determining the sample input length of the training process in the current training round based on the inference result of the target optimization model in the current training round; and performing clustering processing on the training sample data of the training instance based on the sample input length of the training process in the current training round; wherein the data loader is loaded from the host memory into the device memory of the computing card by the processor of the target device.
[0061] The average number of tokens per sample can be the average number of tokens included in multiple training samples for each inference instance. The original training data can be the raw, unprocessed training data. The original training sample data can be the input data determined for each inference instance. The data loader, handled by the processor, loads data from host memory into the compute card's device memory. After data length clustering, it loads batch samples into the device memory for each inference and training instance's compute card. The batch sample loading time is limited by the host PCIe (peripheral component interconnect express, a high-speed serial computer extension bus standard) bandwidth. Inference sample data can be the input data required to be fed into the target optimization model during the inference process of a separate inference and training pipeline. Training sample data can be the input data required to be fed into the target optimization model during the training process of a separate inference and training pipeline.
[0062] Specifically, the data balancing preprocessing of training data can be divided into a coarse data sorting process and a fine data sorting process. Figure 3 This is a schematic diagram of a data coarse sorting process provided by an embodiment of the present invention. In a specific example, such as... Figure 3 As shown, when performing coarse ranking on the original training data of the target optimization model, we can first determine the average number of tokens per inference instance in the current training round of the target optimization model, based on the original training data. This means determining the average prompt word length for each inference instance in the current training round. Assume the original training data contains k samples, and the number of tokens in the i-th sample is... Then the total number of tokens for k sample data is Then the average number of tokens per sample in the original training data is Assuming the input sample batch for each inference instance is s, then the average number of tokens per s samples in each inference instance is... That is, the average number of tokens per inference instance is .
[0063] After determining the average number of tokens per sample for each inference instance, the original training sample data loaded by the data loader of the target device corresponding to the inference instance in the computing card can be further determined based on the average number of tokens per sample for each inference instance. This ensures that the data length of the original training sample data loaded by the data loader of each inference instance in the computing card is basically the same, thereby achieving clustering processing of the original training sample data of the inference instance and thus achieving load balancing operation on the target device corresponding to the inference instance.
[0064] Optionally, the target token size of the input sample for each inference instance can be [size to be specified]. Based on the average number of tokens per inference instance, the iteration order of the data loader for the inference instances can be rearranged so that the total number of tokens in the same inference batch is roughly close to the target requirement, thereby achieving computational load balancing during the inference phase.
[0065] Figure 4 This is a schematic diagram of a data sorting process provided by an embodiment of the present invention. In a specific example, such as... Figure 4 As shown, the original training data of the target optimization model in the current training round can be fine-ranked based on the coarse-ranking results in the current training round. Specifically, after the target optimization model completes the inference process of the current round based on the coarse-ranking results, the inference result executed by the target optimization model based on the coarse-ranking results of the data in the current training round is obtained. The inference result of the target optimization model in the current training round is then concatenated with the input data of the target optimization model in the inference process of the current training round. Based on the concatenation result, the sample input length of the training process in the current training round is determined. Therefore, the training sample data of the training instances can be clustered based on the sample input length of the training process in the current training round.
[0066] Specifically, based on the inference results performed by the target optimization model according to the coarse ranking results of the data in the current training round, the length of the input data prompt and the length of the inference result response of the target optimization model in the inference process of the current training round can be statistically analyzed to obtain the sample input length of the training sample data. When clustering the training sample data of the training instances according to the sample input length of the training process in the current training round, the training sample data with similar sample input lengths can be used as the training sample data of each training instance in the current training round.
[0067] It's important to note that since each sample generates G distinct outputs, reward calculation requires normalization based on these different outputs. During data ranking, the clustering process of the training sample data distributes different candidate outputs from the same training sample data into different mini-batches, necessitating additional cross-device communication to calculate the complete reward. Placing all candidate outputs of the same sample in the same mini-batch can lead to imbalanced training sample data lengths, increasing data padding and computational overhead. Therefore, a threshold can be selected to dynamically adjust the way candidate outputs are distributed across mini-batches.
[0068] It should also be noted that the above data balancing preprocessing process is applicable not only to raw text-based training data, but also to multimodal raw training data, such as raw training data containing "text + image + speech". For non-text raw training data, it can be converted into text tokens through image encoding and speech encoding, and then the total number of tokens can be counted according to the number of text tokens.
[0069] The above technical solution, by performing data balancing preprocessing on the training data of the target optimization model, can ensure that the time consumption of each batch returned by the data loader meets a fixed distribution during the training and inference stages, and that the sample lengths in the same batch are basically consistent, thereby achieving load balancing and providing a foundation for subsequent pipeline optimization.
[0070] S220. Determine the model time information of the target optimization model based on the model parameters of the target optimization model, the data preprocessing information, and the device associated hardware information.
[0071] Accordingly, after performing data equalization preprocessing on the training data of the target optimization model, the model execution time of the target optimization model can be determined based on the data preprocessing information obtained after data equalization preprocessing, combined with the model parameters of the target optimization model and the associated hardware information of the device. It can be understood that the data preprocessing information obtained after data equalization preprocessing can include the length of the inference sample data for inference instances and the length of the training sample data for training instances, etc.
[0072] In an optional embodiment of the present invention, determining the model time information of the target optimization model based on the model parameters of the target optimization model, the data preprocessing information, and the device-associated hardware information may include: determining preprocessed training data based on the data preprocessing information, and loading the preprocessed training data into the data loader of the target device; wherein the preprocessed training data includes preprocessed inference sample data and preprocessed training sample data; running the target optimization model according to the model parameters of the target optimization model, and performing an inference process on the target optimization model based on the preprocessed inference sample data, performing training sample data preprocessing based on the inference results of the inference process, and performing a training process on the target optimization model based on the preprocessed training sample data; calculating the time required for the target device to perform the inference process to obtain the model inference time; and calculating the time required for the target device to perform the training process to obtain the model training time.
[0073] Specifically, preprocessed inference sample data can be the inference sample data required to execute one inference process, determined based on data preprocessing information. Preprocessed training sample data can be the training sample data required to execute one training process, determined based on data preprocessing information.
[0074] In this embodiment of the invention, various methods can be used to determine the model time information of the target optimization model. Optionally, after obtaining the data preprocessing information for data balancing preprocessing, the preprocessed inference sample data for one inference process and the preprocessed training sample data for one training process can be determined based on the data preprocessing information. Further, the preprocessed inference sample data can first be loaded into the data loader of the target device corresponding to the inference instance, and the target optimization model can be run according to the model parameters of the target optimization model to perform one inference process. After the inference process is completed, the training sample data preprocessing process is performed based on the inference result of the inference process, that is, the inference result and the preprocessed inference sample data are concatenated, and each concatenated data is further preprocessed to obtain the final preprocessed training sample data, so that the length of the preprocessed training sample data of the current training instance is kept as consistent as possible, and then the target optimization model is trained once based on the preprocessed training sample data. In this way, by actually performing one inference process and one training process on the target optimization model, the model inference time can be obtained by real-time statistics of the time required for the target device to perform the inference process, and the model training time can be obtained by real-time statistics of the time required for the target device to perform the training process.
[0075] In an optional embodiment of the present invention, the device-associated hardware information includes the peak computing power of the computing card and the peak bandwidth of the memory; the data preprocessing information includes the sample data length, the sample data batch size, and the response token length of each input sample data; determining the model time information of the target optimization model based on the model parameters of the target optimization model, the data preprocessing information, and the device-associated hardware information may include: calculating the time consumption of the first inference stage based on the sample data batch size, the sample data length, and the number of model parameters of the target optimization model; calculating the time consumption of the second inference stage based on the response token length of each input sample data in the second inference stage, the amount of KV (Key Value) cache data, the number of model parameters of the target optimization model, and the peak bandwidth of the memory; calculating the model inference time based on the time consumption of the first inference stage and the time consumption of the second inference stage; calculating the training forward computing power based on the sample data batch size, the sample data length, the response token length of each input sample data, and the number of model parameters of the target optimization model; and calculating the model training time for the current sample based on the training forward computing power and the peak computing power of the computing card.
[0076] The time consumed in the first inference stage can be estimated as the time consumed in the prefill stage (also known as the prefill stage) of the inference process of the target optimization model. The prefill stage is the initial stage of model inference, used to process the context of the input and lay the foundation for subsequent output generation. The time consumed in the second inference stage can be estimated as the time consumed in the recursive prediction stage (also known as the decoding stage) of the generated text in the inference process of the target optimization model. The recursive prediction stage of the model inference process generates subsequent tokens step by step through a self-attention mechanism. The KV cache data size can be the size of the KV cache (an optimization technique used to accelerate autoregressive generation tasks such as text generation) that caches the number of tokens in the batch of inference instance samples of all layers of the target optimization model.
[0077] For example, the estimation process of model latency information is illustrated using a single-card setup. For multi-card setups, the latency caused by communication between multiple cards needs to be considered when estimating model latency information. Assume that FP16 is used for inference training of the target optimization model, and that the peak computing power of the target device's FP16 computing card is P, and the peak bandwidth of HBM (High Bandwidth Memory) is [missing value]. For a batch size of B and a length of S, the target optimization model can generate a response of length C for each sample. The inference process can be divided into a prefill phase and a recursive prediction phase (decoding). The prefill phase is computationally intensive, while the decoding phase is I / O (Input / Output) intensive. For the prefill phase, it is assumed that its time consumption is the same as that of the first inference phase. The first reasoning stage takes time The optimal model can be optimized based on the sample data batch size B, sample data length S, and the number of model parameters. Calculated, i.e. During the decoding phase, generating the first token requires reading the KV cache of S tokens, the second requires reading S+1 tokens, and the Cth requires reading S+C-1 tokens. Assuming the total KV cache data size for loading all S tokens from all layers is... and will As for the amount of KV cached data, the time consumed in the second inference stage... The time consumed in the second inference stage consists of KV cache loading time and model weight loading time. This can be calculated based on the response token length, KV cache data volume, model parameter count of the target optimization model, and peak memory bandwidth for each input sample data in the second inference stage. Here, BPE represents the number of bytes occupied by a single element of the model parameter in the target optimization model. Therefore, the model inference time can be calculated based on the time consumed in the first inference stage and the time consumed in the second inference stage. That is, the model inference time is .
[0078] Since the model training process is a computationally intensive task, assuming the forward computation required for one training iteration is... Then, based on the sample data batch size B, sample data length S, response token length C for each input sample data, and the number of model parameters of the target optimization model, we can determine the optimal model. Calculate the forward computing power of training ,Right now The computational power required for one reverse training iteration is approximately 2. Therefore, the training time of the model for the current sample can be calculated based on the forward computing power and the peak computing power of the computing card. ,Right now .
[0079] S230. Determine the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time.
[0080] S240. Determine the number of first target devices required for inference and the number of second target devices required for training of the target optimization model based on the target ratio of training instances to inference instances of the target optimization model.
[0081] The target devices of the first target number and the target devices of the second target number are used to complete the asynchronous pipelined training process of the push-train separation of the target optimization model.
[0082] In an optional embodiment of the present invention, determining the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time may include: calculating the inference instance time of the number of parallel execution samples based on the number of parallel execution samples of the target optimization model, the number of inference instances, the number of batch samples of inference instances, the number of output data generated per sample, and the model inference time; calculating the training instance time of the number of parallel execution samples based on the number of parallel execution samples of the target optimization model, the number of batch samples of training instances, the number of training instances, and the model training time; and determining the target ratio of training instances to inference instances of the target optimization model based on a first constraint relationship between the inference instance time and the training instance time.
[0083] Here, the number of parallel execution samples can be the number of samples in one iteration cycle of the target optimization model when performing the training task. The inference instance time can be the time taken for an inference instance to perform one inference process. The training instance time can be the time taken for a training instance to perform one training process. The first constraint relationship can be the constraint relationship set for the inference instance time and the training instance time when solving for the target ratio.
[0084] Specifically, calculating the inference instance time for the number of parallel execution samples based on the target optimization model's number of parallel execution samples, number of inference instances, number of batch samples for inference instances, number of output data generated per sample, and model inference time can include: calculating the inference instance time for the number of parallel execution samples based on the following formula:
[0085]
[0086] Where W represents the number of parallel execution samples of the target optimization model, m represents the number of inference instances, s represents the number of batch samples of inference instances (i.e., the number of inference samples), and G represents the number of output data generated per sample (i.e., the number of responses generated per sample). This indicates the time taken for model inference.
[0087] Specifically, calculating the training instance time for the number of parallel execution samples based on the number of parallel execution samples of the target optimization model, the number of batch samples of training instances, the number of training instances, and the model training time can include: calculating the training instance time for the number of parallel execution samples based on the following formula:
[0088]
[0089] Where n represents the number of training instances, and b represents the number of batch samples for each training instance, i.e., the mini-batch size averaged across all training instances. This indicates the time taken to train the model.
[0090] Furthermore, based on the first constraint relationship between the inference instance time and the training instance time, determining the target ratio of training instances to inference instances for the target optimization model can include: determining the first constraint relationship between the inference instance time and the training instance time based on the following formula:
[0091]
[0092] In other words, the above formula means: make the time spent on inference instances and the time spent on training instances as similar as possible. It should be noted that the time spent on inference instances here can be the time spent on inference instances in the current training round, while the time spent on training instances can be the time spent on training instances that need to use the inference results of the inference instances in the current training round. That is, try to make the time spent on inference instances in the next training round synchronized with the time spent on training instances in the current training round, and complete them as synchronously as possible, so as to achieve the parallel execution of inference and training.
[0093] Furthermore, the target ratio of training instances to inference instances for the target optimization model is determined based on the following formula:
[0094]
[0095] That is, assuming there are n training instances and m inference instances, for a given inference instance, it infers s samples, and each sample generates G responses. The time required to generate a total of sG responses is... For a training example, the time taken to perform one forward (or two, depending on whether there is a reference model) and backpropagation for b samples is... To allow pipeline overlap, training instances and inference instances should take as little time as possible to process the same samples. Suppose there are W samples to process in parallel; then the target ratio can be expressed as: In other words, the optimal ratio of training instances to inference instances (n / m) can be calculated using the number of generated samples, batch size, and the ratio of training to inference time. and It can be obtained through actual measurement or through theoretical modeling. The above method for determining the target ratio is applicable when the length of the response output by the inference result is relatively stable during the training process of the target optimization model.
[0096] In an optional embodiment of the present invention, determining the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time may include: calculating the total target switching time based on the switching time of the target device; calculating the idle time of the computing card for training the number of parallel execution samples of the target optimization model without switching training and inference instances based on the model training time and the model inference time; calculating the target switching time of training instances and inference instances based on the total target switching time, the idle time of the computing card for training the number of parallel execution samples of the target optimization model without switching training and inference instances, and the idle time of the computing card for training the number of parallel execution samples of the target optimization model with switching training and inference instances; and determining the target ratio of training instances to inference instances of the target optimization model based on the target switching time of training instances and inference instances and the second constraint relationship.
[0097] The total target switching time can be defined as the total switching time between the inference instance and the training instance at regular intervals over a certain number of training rounds. The target switching time can be defined as the final switching time calculated for the training instance and the inference instance. The second constraint can be defined as the constraint that minimizes the target switching time.
[0098] In the asynchronous pipelined training of a target optimization model, where training and inference are separated, complex situations may arise with each training epoch. For example, the length of the model's output may change during asynchronous pipelined training, leading to variations in the length of the input data for training. Therefore, the target ratio of the specific training instances to inference instances for each training epoch needs to be dynamically adjusted. Dynamically adjusting the target ratio requires balancing the performance gains from switching between inference and training instances against the time cost of the switch itself. In some target optimization model training frameworks, the length of the response obtained by the inference instance gradually increases as training progresses, causing the training and inference times to become variable and increase over time.
[0099] Assuming the number of training steps t and the average response length ( The change of ) roughly satisfies linearity, that is ,in These are all coefficients, and their values can be determined through fitting. The inference process of the target optimization model mainly consists of a prefill stage and a decoding stage. Increasing the length of the inference response leads to a sharp increase in the time consumed in the decoding stage. However, since the input prompt remains unchanged, the prefill stage is largely unaffected. Generally, in the decoding stage, the performance bottleneck occurs due to memory access speed limitations during model execution, and the inference time is linearly related to the length of the inference result. During the training process, a forward propagation is required for the concatenated prompt and response. During the forward propagation in training, most of the time spent on operator execution is consumed in arithmetic calculations rather than data transmission. Furthermore, the forward propagation time is quadratically related to the length of the input sequence, meaning... , As training progresses, the average training and inference times will change. Switching between training and inference instances requires a trade-off between pipeline evacuation and switching overhead. This switching overhead includes the following: adding an inference instance requires unloading the memory-intensive components such as the training model, optimizer, and gradients, then loading the inference model, compiling the computation graph, and repartitioning the key-value cache; conversely, adding a training instance requires unloading the inference model and its key-value cache, reloading the training model, and restoring the optimizer and gradient states. Furthermore, if gradient splitting or model splitting strategies are used during training, the splitting parameters must be redistributed and adjusted each time a training instance is added or removed.
[0100] Figure 5 This is a flowchart illustrating a process for dynamically determining the target ratio of training instances and inference instances, provided by an embodiment of the present invention. In a specific example, such as... Figure 5 As shown, assume that the time taken for one switch between inference and training instances is... and every The switching occurs once every training epoch, with the aim of minimizing compute card idle time and switching time. This needs to be determined... The optimal solution. The above problem can be defined as a mathematical optimization problem. That is, assuming that in In each training round, every Perform a switch between inference and training instances. Taking a constant step size as an example, calculate the total target switching time. It can be:
[0101]
[0102] in, Indicates in In each training step, every... The total handover time for one handover. This indicates the total number of training rounds. This represents the interval between the number of steps required to switch between a training instance and an inference instance. This represents the time required to switch between an inference instance and a training entity.
[0103] Specifically, the idle time of the computing card for the number of parallel execution samples of the training target optimization model can be calculated based on the model training time and model inference time without switching training and inference instances. First, the time saved is the reduction in the idle time of the computing card due to the switching, as shown in the following formula:
[0104]
[0105] The aforementioned idle time of the computation card is the computation time for training W sample data. Therefore, for the total amount of data in training round t... Without switching between training and inference instances, the calculation of the number of parallel execution samples for optimizing the training objective model and the idle time are calculated. for:
[0106]
[0107] It should be noted that, Figure 5 In the illustrated process, the idle time is first calculated to determine the number of parallel execution samples for optimizing the training objective model without switching training inference instances. Then calculate the total target switching time. However, the embodiments of the present invention do not apply to and The order of calculation is restricted. That is, calculations can be performed first. Recalculate You can also calculate first. Recalculate Or it can be calculated simultaneously. and .
[0108] Assuming the target optimization model is switched, for Switch to During the period, the ratio of balanced inference instances to training instances is: and The corresponding time consumption is and The idle time of the computation card for the i-th training instance and inference instance switching is . The time saved by the idle time of the computing card after the ratio switch, that is, the idle time of the computing card for the number of parallel execution samples of the training objective optimization model under the condition of training inference instance switching, is:
[0109]
[0110] Adding the time consumed by multiple switching, based on the total target switching time, the idle time of the computing card for the number of parallel execution samples of the target optimization model without training and inference instance switching, and the idle time of the computing card for the number of parallel execution samples of the target optimization model with training and inference instance switching, the target switching time for training instances and inference instances is calculated as follows:
[0111]
[0112] Furthermore, an objective function can be constructed based on the target switching time of training instances and inference instances. The formula is as follows:
[0113]
[0114] Based on the target switching time of training and inference instances, an objective function is constructed, which, combined with the second constraint, leads to the following problem:
[0115]
[0116]
[0117] in, This represents the second constraint relation for minimizing the objective function. Solving the optimal solution to the above problem determines the target ratio of training instances to inference instances in the current optimization model. If the target ratio of training instances to inference instances calculated in the current round has changed compared to the previous round, the current real-time calculated ratio can be used to improve training efficiency. If the target ratio of training instances to inference instances calculated in the current round has not changed compared to the previous round, the ratio from the previous round can be kept unchanged. This allows for periodic switching of instance allocation, thereby improving the training efficiency of the objective optimization model in real time.
[0118] It should be noted that the above-mentioned dynamic calculation of the target ratio is applicable to dynamic switching scenarios where the length of the inference response changes continuously during the training of the target optimization model. It is also applicable to scenarios where the model changes and / or the cluster computing card resources change. If a faulty node occurs in the target device, it is necessary to recalculate and configure the ratio of inference instances and training instances.
[0119] After determining the target ratio n / m between training instances and inference instances of the target optimization model, the number of target devices required for inference and training of the target optimization model can be determined based on the target ratio n / m. Then, the overall training process of separating inference and training asynchronous pipeline can be started for the target optimization model based on the determined number of instance devices.
[0120] Figure 6 This is a schematic diagram illustrating the overall training process of an asynchronous pipelined inference and training instance for a target optimization model, provided by an embodiment of the present invention. In a specific example, such as... Figure 6 As shown, GPU0-GPU3 represent inference cards with 4 inference instances, and GPU4-GPU9 represent training cards with 6 training instances. Figure 6 The numerical identifier in the diagram represents a batch. For example, during inference, GPU0 trains samples from three batches (batch 1, batch 2, and batch 3) in the first training epoch. GPUs 0 through 3 can train a total of 12 batches per training epoch, and a total of 48 batches after the entire training process. Similarly, in the training process, GPU4 trains samples from two batches (batch 1 and batch 2) in the first training epoch, GPU5 trains samples from two batches (batch 3 and batch 4) in the first training epoch, and so on. GPUs 4 through 9 can train a total of 12 batches per training epoch, and a total of 48 batches after the entire training process.
[0121] like Figure 6As shown, firstly, all inference instances complete the inference of the first batch. For example, GPU0 infers based on batches 1-3, GPU3 infers based on batches 10-12, and so on. At this point, since there are no inference results yet, the training instances are idle. Because each inference instance's batch has undergone data balancing preprocessing, the inference time for each inference card is basically the same during the first round of inference for GPUs 0-GPU 3, and the inference cards can basically complete the first round of inference synchronously. After the inference instances complete the inference of batch 1 in the first round, they immediately begin the inference of the second batch, i.e., batches 13-24. At the same time, all training instances can use the inference results of the first batch combined with the first batch to start training. That is, GPU4 concatenates the inference results of batch 1 to train batch 1, and concatenates the inference results of batch 2 to train batch 2; GPU7 concatenates the inference results of batch 7 to train batch 7, and concatenates the inference results of batch 8 to train batch 8, and so on. Because each training instance's batch undergoes data balancing preprocessing, the training time for each GPU (GPU4-GPU9) is essentially the same during the first round of training, allowing them to complete the first round of training almost synchronously. This continues, meaning the entire model training pipeline is essentially at full saturation, with no significant periods of idle computing resources. After the inference instance completes the inference of the last batch, the training instances will continue training the penultimate batch. Finally, while the training instances are processing the last batch, all inference instances enter an idle state. In other words, except for the very first and last batches which may be idle, the entire pipeline is saturated for the rest of the time.
[0122] The above-mentioned asynchronous pipelined training process, which separates training and inference, constructs an asynchronous pipeline mechanism. By covering the time the training card waits for the inference card to synchronize data, and by utilizing the fixed distribution characteristics after preprocessing to reduce the idle time of the pipeline's computing cards, the overall training time can be significantly reduced.
[0123] It should be noted that, although Figure 6 This illustrates an ideal scenario where the inference time of one round is the same as the training time of one round. In reality, during asynchronous pipelined training with separate inference and training, the inference time of inference instances is usually greater than the training time of training instances. Because each round of training instances must wait for the inference process of the corresponding round of inference instances to finish before it can begin, there may be brief waiting times between training rounds. Figure 6As shown, in actual practice, there may be a short period of idle waiting time between batch2 and batch13, but this time has little impact on the overall time consumption and can be basically ignored.
[0124] Figure 7 This is a schematic diagram comparing the time consumption of different push-train coupling training methods provided in an embodiment of the present invention. In a specific example, such as... Figure 7 As shown, the training execution speed of the entire model optimized by inference instances in this embodiment of the invention is significantly better than the original push-train coupled training mode, achieving efficient utilization of computing resources and providing a better solution for large-scale model training. Although the optimized scheme has idle periods at the beginning and end, its overall time is still significantly less than the unoptimized push-train coupled version (even though the latter has no idle periods). The push-train separated synchronous pipeline version takes the longest time because most of the time is spent waiting for synchronization operations.
[0125] This invention provides a method for accelerating model training by separating inference and training into asynchronous pipelines. Through improvements in data preprocessing, cluster resource allocation optimization, and asynchronous pipelines, the method optimizes the configuration of computing card resources based on the architecture parameters of the computing card. By using cross-pipeline technology for training and inference, the method avoids waiting for computing card resources, improves the utilization of hardware resources during model training, reduces the time spent on model training, and improves the efficiency of model training.
[0126] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions.
[0127] It should be noted that any arrangement or combination of the technical features in the above embodiments also falls within the protection scope of this invention.
[0128] Figure 8 This is a schematic diagram of a model training optimization device provided in an embodiment of the present invention, such as... Figure 8 As shown, the device includes: a model time acquisition module 310, an inference-training instance ratio determination module 320, and a target device quantity determination module 330, wherein:
[0129] The model time consumption information acquisition module 310 is used to acquire the model time consumption information of the target optimization model through the device-associated hardware information of the target device running the target optimization model; wherein, the model time consumption information includes the model training time and model inference time of the target optimization model;
[0130] The inference-training instance ratio determination module 320 is used to determine the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time.
[0131] The target device quantity determination module 330 is used to determine the first target device quantity required for inference and the second target device quantity required for training of the target optimization model based on the target ratio of training instances to inference instances of the target optimization model.
[0132] The target devices of the first target number and the target devices of the second target number are used to complete the asynchronous pipelined training process of the push-train separation of the target optimization model.
[0133] This invention obtains model time information, such as model training time and model inference time, from the device-associated hardware information of the target device running the target optimization model. Then, based on the model training time and model inference time, it determines the target ratio of training instances to inference instances of the target optimization model. Based on this ratio, it determines the number of first target devices required for inference and the number of second target devices required for training. After determining these numbers, the asynchronous pipelined training process of the target optimization model can be completed using the first and second target devices. This technical solution optimizes the resource allocation of the target device during model inference and training by using the device-associated hardware information of the target device. This significantly reduces the waiting time for computing card resources on the target device, solving the problems of low hardware resource utilization and low overall training efficiency in existing models using a push-train coupling training method. It improves the hardware resource utilization of the model during push-train coupling training, reduces the latency of push-train coupling training, and improves the efficiency of push-train coupling training.
[0134] Optionally, the model time acquisition module 310 is further configured to: perform data equalization preprocessing on the training data of the target optimization model in each training round of the target optimization model to obtain data preprocessing information; and determine the model time information of the target optimization model based on the model parameters of the target optimization model, the data preprocessing information, and the device-associated hardware information.
[0135] Optionally, the model time consumption information acquisition module 310 is further configured to: determine the average number of tokens per inference instance based on the original training data of the target optimization model in the current training round; determine the original training sample data loaded by the data loader of the inference instance in the computing card based on the average number of tokens per inference instance, so as to achieve clustering processing of the inference sample data of the inference instance; determine the sample input length of the training process in the current training round based on the inference result of the target optimization model in the current training round; and perform clustering processing on the training sample data of the training instance based on the sample input length of the training process in the current training round; wherein the data loader is loaded from the host memory into the device memory of the computing card by the processor of the target device.
[0136] Optionally, the model time acquisition module 310 is further configured to: determine preprocessed training data based on the data preprocessing information, and load the preprocessed training data into the data loader of the target device; wherein the preprocessed training data includes preprocessed inference sample data and preprocessed training sample data; run the target optimization model according to the model parameters of the target optimization model, and perform an inference process on the target optimization model according to the preprocessed inference sample data, perform training sample data preprocessing according to the inference results of the inference process, and perform a training process on the target optimization model according to the preprocessed training sample data; calculate the time required for the target device to perform the inference process to obtain the model inference time; calculate the time required for the target device to perform the training process to obtain the model training time.
[0137] Optionally, the device-associated hardware information includes the peak computing power of the computing card and the peak bandwidth of the memory; the data preprocessing information includes the sample data length, the sample data batch size, and the response token length for each input sample data; the model time acquisition module 310 is further configured to: calculate the time consumption of the first inference stage based on the sample data batch size, the sample data length, and the number of model parameters of the target optimization model; calculate the time consumption of the second inference stage based on the response token length of each input sample data in the second inference stage, the amount of KV cache data, the number of model parameters of the target optimization model, and the peak bandwidth of the memory; calculate the model inference time based on the time consumption of the first inference stage and the time consumption of the second inference stage; calculate the training forward computing power based on the sample data batch size, the sample data length, the response token length of each input sample data, and the number of model parameters of the target optimization model; and calculate the model training time for the current sample based on the training forward computing power and the peak computing power of the computing card.
[0138] Optionally, the inference-training instance ratio determination module 320 is further configured to: calculate the inference instance time of the number of parallel execution samples based on the number of parallel execution samples of the target optimization model, the number of inference instances, the number of batch samples of inference instances, the number of output data generated by each sample, and the model inference time; calculate the training instance time of the number of parallel execution samples based on the number of parallel execution samples of the target optimization model, the number of batch samples of training instances, the number of training instances, and the model training time; and determine the target ratio of training instances to inference instances of the target optimization model based on the first constraint relationship between the inference instance time and the training instance time.
[0139] Optionally, the inference-training instance ratio determination module 320 is further configured to: calculate the total target switching time based on the switching time of the target device; calculate the idle time of the computing card for training the number of parallel execution samples of the target optimization model without switching training and inference instances based on the model training time and the model inference time; calculate the target switching time of training instances and inference instances based on the total target switching time, the idle time of the computing card for training the number of parallel execution samples of the target optimization model without switching training and inference instances, and the idle time of the computing card for training the number of parallel execution samples of the target optimization model with switching training and inference instances; and determine the target ratio of training instances and inference instances of the target optimization model based on the target switching time of training instances and inference instances and the second constraint relationship.
[0140] Optionally, the asynchronous pipelined training process of the target optimization model is an interleaved execution process of inference and training; the first process in the asynchronous pipelined training process of the target optimization model is the inference process, and there is no inference waiting time between each inference process.
[0141] The above-described model training optimization apparatus can execute the model training optimization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the model training optimization method provided in any embodiment of the present invention.
[0142] Figure 9A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0143] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0144] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0145] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as model training optimization methods.
[0146] Optionally, the model training optimization method may include: obtaining model time information of the target optimization model through device-associated hardware information of the target device running the target optimization model; wherein the model time information includes model training time and model inference time of the target optimization model; determining a target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time; determining a first number of target devices required for inference and a second number of target devices required for training of the target optimization model based on the target ratio of training instances to inference instances of the target optimization model; wherein the target devices of the first number of target devices and the target devices of the second number of target devices are used to complete the asynchronous pipelined training process of the target optimization model.
[0147] In some embodiments, the model training optimization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the model training optimization method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the model training optimization method by any other suitable means (e.g., by means of firmware).
[0148] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0149] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0150] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0152] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0153] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0154] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0155] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A model training optimization method, characterized in that, include: The model time information of the target optimization model is obtained by using the device-associated hardware information of the target device running the target optimization model; wherein, the model time information includes the model training time and model inference time of the target optimization model; the device-associated hardware information includes the peak computing power of the computing card and the peak memory bandwidth; The target ratio of training instances to inference instances of the target optimization model is determined based on the model training time and the model inference time. The number of first target devices required for inference and the number of second target devices required for training of the target optimization model are determined based on the target ratio of training instances to inference instances of the target optimization model. The target devices of the first target number and the target devices of the second target number are used to complete the asynchronous pipelined training process of the push-train separation of the target optimization model.
2. The method according to claim 1, characterized in that, The step of obtaining the model time information of the target optimization model by associating hardware information with the device running the target optimization model includes: In each training round of the target optimization model, the training data of the target optimization model is subjected to data equalization preprocessing to obtain data preprocessing information; The model time information of the target optimization model is determined based on the model parameters of the target optimization model, the data preprocessing information, and the device-associated hardware information.
3. The method according to claim 2, characterized in that, The data balancing preprocessing of the training data for the target optimization model includes: In the current training round of the target optimization model, the average number of tokens per inference instance is determined based on the original training data of the target optimization model. The original training sample data loaded into the computing card by the data loader of the inference instance is determined based on the average number of tokens in each inference instance, so as to realize the clustering processing of the inference sample data of the inference instance. The sample input length of the training process in the current training round is determined based on the inference results of the current training round of the target optimization model. The training sample data of the training instance is clustered based on the sample input length of the training process in the current training round. The data loader is loaded from the host memory into the device memory of the computing card by the processor of the target device.
4. The method according to claim 3, characterized in that, The step of determining the model time information of the target optimization model based on the model parameters of the target optimization model, the data preprocessing information, and the device-associated hardware information includes: Preprocessing training data is determined based on the data preprocessing information, and the preprocessing training data is loaded into the data loader of the target device; wherein, the preprocessing training data includes preprocessing inference sample data and preprocessing training sample data; The target optimization model is run according to the model parameters of the target optimization model, and an inference process is performed on the target optimization model according to the preprocessed inference sample data. The training sample data is preprocessed according to the inference results of the inference process, and a training process is performed on the target optimization model according to the preprocessed training sample data. The time required for the target device to execute the inference process is calculated to obtain the model inference time. The training time of the model is obtained by calculating the time required for the target device to execute the training process.
5. The method according to claim 3, characterized in that, The data preprocessing information includes the sample data length, sample data batch size, and response token length for each input sample data; the step of determining the model time information of the target optimization model based on the model parameters of the target optimization model, the data preprocessing information, and the device-associated hardware information includes: The time consumed in the first inference stage is calculated based on the sample data batch size, the sample data length, and the number of model parameters of the target optimization model. The time consumed in the second inference stage is calculated based on the response token length of each input sample data in the second inference stage, the amount of key-value cache data, the number of model parameters of the target optimization model, and the peak memory bandwidth. The inference time of the model is calculated based on the time consumed in the first inference stage and the time consumed in the second inference stage. The training forward computing power is calculated based on the sample data batch size, the sample data length, the response token length of each input sample data, and the number of model parameters of the target optimization model. The training time of the model for the current sample is calculated based on the forward computing power of the training and the peak computing power of the computing card.
6. The method according to claim 4 or 5, characterized in that, Determining the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time includes: The inference instance time of the number of parallel execution samples of the target optimization model is calculated based on the number of inference instances, the number of batch samples of the inference instances, the number of output data generated by each sample, and the inference time of the model. The training instance time for the number of parallel execution samples is calculated based on the number of parallel execution samples of the target optimization model, the number of training instance batch samples, the number of training instances, and the training time of the model. Based on the first constraint relationship between the inference instance time consumption and the training instance time consumption, the target ratio of the training instance to the inference instance of the target optimization model is determined.
7. The method according to claim 4 or 5, characterized in that, Determining the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time includes: Calculate the total target switching time based on the switching time of the target device; The idle time of the calculation card is used to calculate the number of parallel execution samples for training the target optimization model without switching training and inference instances, based on the model training time and the model inference time. The target switching time for training instances and inference instances is calculated based on the total target switching time, the idle time of the computing card for the number of parallel execution samples of the target optimization model when training and inference instances are not switched, and the idle time of the computing card for the number of parallel execution samples of the target optimization model when training and inference instances are switched. Based on the target switching time and the second constraint relationship between the training instance and the inference instance, the target ratio of the training instance and the inference instance of the target optimization model is determined.
8. The method according to any one of claims 6 or 7, characterized in that, The asynchronous pipelined training process of the target optimization model is an interleaved execution process of inference and training; the first process in the asynchronous pipelined training process of the target optimization model is the inference process, and there is no inference waiting time between each inference process.
9. A model training optimization device, characterized in that, include: The model time consumption information acquisition module is used to acquire the model time consumption information of the target optimization model through the device-associated hardware information of the target device running the target optimization model; wherein, the model time consumption information includes the model training time and model inference time of the target optimization model; the device-associated hardware information includes the peak computing power of the computing card and the peak bandwidth of the memory. The inference-training instance ratio determination module is used to determine the target ratio of training instances to inference instances of the target optimization model based on the model training time and the model inference time. The target device quantity determination module is used to determine the first target device quantity required for inference and the second target device quantity required for training of the target optimization model based on the target ratio of training instances to inference instances of the target optimization model. The target devices of the first target number and the target devices of the second target number are used to complete the asynchronous pipelined training process of the push-train separation of the target optimization model.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that is executed by the at least one processor, such that the at least one processor is able to perform the model training optimization method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the model training optimization method according to any one of claims 1-8.
12. A computer program product comprising a computer program / instructions, wherein, When the computer program / instructions are executed by the processor, they implement the model training optimization method according to any one of claims 1-8.
Citation Information
Patent Citations
Model training deployment method and electronic equipment
CN118277040A
Service processing method and apparatus, and computer device, storage medium and program product
WO2025082020A1