DeepSpeed-based gradient unloading optimization method and device and storage medium
By unloading FP16 gradient data to the expansion card in real time during the gradient accumulation process of DeepSpeed, the problem of excessive video memory usage in large-scale deep learning model training is solved, and the effect of significantly reducing GPU memory usage is achieved and training efficiency is improved.
Patent Information
- Application Number
- CN202510276483.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
AI Technical Summary
When DeepSpeed conducts large-scale deep learning model training, the video memory occupies too much, making it difficult to meet the needs of data storage and computing.
During the gradient accumulation process, the FP16 gradient data is unloaded into the expansion card in real time, and the gradient unloading optimization method based on DeepSpeed is adopted, including obtaining the current FP16 gradient data, reading the last saved FP16 gradient data from the expansion card, accumulating the two, and rewrite the accumulated gradient data back to the expansion card.
It significantly reduces GPU memory footprint, improves the training efficiency of large-scale models, and solves the problem of excessive video memory footprint.
Smart Images

Figure CN120144304A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model training, and particularly to a gradient offloading optimization method, device, and storage medium based on DeepSpeed. Background Art
[0002] DeepSpeed is a deep learning optimization library designed to optimize memory for large-scale deep learning model training. During model training, DeepSpeed adopts a strategy of offloading gradient data in the video memory to the memory, and performs gradient accumulation operations by maintaining FP16 gradients in the memory. However, as the scale of deep learning models continues to expand, the number of parameters involved in model training has become extremely large. In this case, neither the capacity of the video memory nor the memory can meet the requirements for data storage and operation during deep learning model training and inference. Summary of the Invention
[0003] The main purpose of this application is to provide a gradient offloading optimization method, device, and storage medium based on DeepSpeed, aiming to solve the technical problem of excessive video memory occupation when DeepSpeed performs large-scale deep learning model training.
[0004] To achieve the above object, an embodiment of this application provides a gradient offloading optimization method based on DeepSpeed. The gradient offloading optimization method based on DeepSpeed includes: In response to performing a gradient accumulation operation, obtain the current FP16 gradient data; Read the previously saved FP16 gradient data from the expansion card; Accumulate the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data; Rewrite the accumulated gradient data back to the expansion card.
[0005] In one embodiment, before the step of in response to performing a gradient accumulation operation, obtain the current FP16 gradient data, it includes: Pre-allocate a first buffer for the FP16 gradient data; Obtain the current FP16 gradient data and copy the current FP16 gradient data to the first buffer; Write the current FP16 gradient data into the expansion card through the first buffer.
[0006] In one embodiment, after the step of obtain the current FP16 gradient data and copy the current FP16 gradient data to the first buffer, it further includes: Determine the partition number, gradient offset, and gradient length of the current FP16 gradient data; According to the partition number, the gradient offset, and the gradient length, write the current FP16 gradient data into the gradient file with the corresponding number through the first buffer, and the gradient file is located in the expansion card.
[0007] In one embodiment, the step of adding the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data includes: Add the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data; Rewrite the accumulated gradient data back to the corresponding gradient file in the expansion card according to the partition number, gradient offset, and gradient length.
[0008] In one embodiment, the step of rewriting the accumulated gradient data back to the expansion card includes: After the gradient accumulation operation is completed, write the accumulated gradient data as the latest FP16 gradient data into the expansion card, and the latest FP16 gradient data will overwrite the previously saved FP16 gradient data and the current FP16 gradient data in the expansion card.
[0009] In one embodiment, after the step of rewriting the accumulated gradient data back to the expansion card, it further includes: When performing optimizer parameter update, read the latest FP16 gradient data from the expansion card and convert the latest FP16 gradient data into FP32 gradient data to complete the optimizer parameter update.
[0010] In one embodiment, before the step of reading the latest FP16 gradient data from the expansion card and converting the latest FP16 gradient data into FP32 gradient data to complete the optimizer parameter update when performing optimizer parameter update, it further includes: Allocate a second buffer and a third buffer. The second buffer is used to store the latest FP16 gradient data read from the expansion card, and the third buffer is used to store the converted FP32 gradient data; Read the latest FP16 gradient data from the expansion card, transfer the latest FP16 gradient data to the host memory through the second buffer, and convert the latest FP16 gradient data into the FP32 gradient data in the host memory; Store the FP32 gradient data in the third buffer to complete the parameter update.
[0011] In one embodiment, the step of allocating a second buffer and a third buffer for the FP16 gradient data and the FP32 gradient data respectively further includes: Determine the sizes of the second buffer and the third buffer according to the number of partitioning parameters. Wherein, the size of the second buffer is the number of partitioning parameters multiplied by 2 bytes, and the size of the third buffer is the number of partitioning parameters multiplied by 4 bytes.
[0012] An embodiment of the present application further provides a gradient offloading optimization device based on DeepSpeed. The gradient offloading optimization device based on DeepSpeed includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the gradient offloading optimization method based on DeepSpeed as described above.
[0013] An embodiment of the present application further provides a storage medium. The storage medium is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the gradient offloading optimization method based on DeepSpeed as described above.
[0014] An embodiment of the present application discloses a gradient offloading optimization method based on DeepSpeed. By responding to the execution of the gradient accumulation operation, obtain the current FP16 gradient data; read the previously saved FP16 gradient data from the expansion card; accumulate the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data; rewrite the accumulated gradient data back to the expansion card. The present application significantly reduces the GPU memory occupancy and improves the training efficiency of large-scale models by offloading gradient data to the expansion card in real time during the gradient accumulation process. Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of the first embodiment of the gradient offloading optimization method based on DeepSpeed involved in the solution of the embodiment of the present application; Figure 2 It is a schematic flowchart of the brief gradient data stream based on DeepSpeed of the first embodiment of the gradient offloading optimization method based on DeepSpeed involved in the solution of the embodiment of the present application; Figure 3 It is a schematic flowchart of the second embodiment of the gradient offloading optimization method based on DeepSpeed involved in the solution of the embodiment of the present application; Figure 4 It is a schematic flowchart of the third embodiment of the gradient offloading optimization method based on DeepSpeed involved in the solution of the embodiment of the present application; Figure 5 This is a schematic diagram of the brief process of the backpropagation stage of the third embodiment of the gradient offloading optimization method based on DeepSpeed involved in the solution of this application example; Figure 6 This is a schematic diagram of the brief process of the optimizer state update stage of the fourth embodiment of the gradient offloading optimization method based on DeepSpeed involved in the solution of this application example; Figure 7 This is a schematic diagram of the structure of the gradient offloading optimization device based on DeepSpeed involved in the solution of this application example.
[0016] The realization of the purpose, functional characteristics and advantages of this application will be further described with reference to the accompanying drawings in combination with the embodiments. Detailed implementation manners
[0017] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0018] DeepSpeed is a deep learning optimization library aimed at achieving memory optimization for large-scale deep learning model training. During model training, DeepSpeed maintains FP16 gradients in memory for gradient accumulation. After completing gradient accumulation, DeepSpeed converts the FP16 gradients into FP32 gradients and offloads them to the expansion card to ensure the calculation accuracy in the optimizer state update stage. However, maintaining FP16 gradients in memory and storing FP32 gradients on the expansion card significantly increases the video memory occupancy and reduces the resource utilization efficiency.
[0019] To solve the above-mentioned defects existing in the related technologies, the embodiment of this application proposes a gradient offloading optimization method based on DeepSpeed. This method obtains the current FP16 gradient data in response to performing the gradient accumulation operation; reads the previously saved FP16 gradient data from the expansion card; accumulates the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data; rewrites the accumulated gradient data back to the expansion card. By offloading the gradient data to the expansion card in real time during the gradient accumulation process, this application significantly reduces the GPU memory occupancy and improves the training efficiency of large-scale models.
[0020] It should be noted that the execution subject of this embodiment can be a DeepSpeed-based gradient offloading optimization device capable of implementing the above functions, or a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, etc., or a virtualization device deployed in a cloud computing platform or a distributed computing environment, which can support the training and optimization of large-scale deep learning models. Hereinafter, a DeepSpeed-based gradient offloading optimization device will be taken as an example to illustrate this embodiment and the following embodiments.
[0021] For the DeepSpeed-based gradient offloading optimization method according to the first embodiment proposed in this application, please refer to Figure 1 , and this method includes steps S10 to S40: Step S10: In response to the execution of the gradient accumulation operation, obtain the current FP16 gradient data.
[0022] In the training of deep learning models, especially when dealing with large-scale models or large batches of data, due to the memory capacity limitation of hardware (such as GPUs), it is impossible to process the entire batch of data at once.
[0023] The gradient accumulation operation refers to dividing a large batch of data into multiple small batches of data, performing forward propagation and backward propagation calculations on these small batches of data in sequence, and then accumulating the gradient data obtained from the calculation of each small batch of data. When the accumulation reaches the preset number of times, the model parameters are updated according to the accumulated gradient data. Through the gradient accumulation operation, the effect of large batch data training can be simulated under limited hardware memory conditions, improving the stability and convergence speed of model training.
[0024] Before model training, the number of gradient accumulation steps is preset. In the backward propagation stage, when the gradient data is generated, the system will first determine whether the gradient accumulation operation needs to be executed. If the current step has reached the preset accumulation step (indicating that the gradient accumulation operation needs to be executed), the current FP16 gradient data can be obtained by accessing the grad attribute of the model parameters and converting it to the FP16 format.
[0025] After each forward propagation and backward propagation of a small batch of data is completed, the gradient data corresponding to this small batch will be obtained. The number of gradient accumulation steps indicates how many small batches of gradient data need to be accumulated.
[0026] Step S20: Read the previously saved FP16 gradient data from the expansion card.
[0027] In this embodiment, an expansion card refers to a device specifically used to expand hardware functions or storage capacity. In gradient offloading optimization, the expansion card can serve as additional space for storing gradient data to relieve the storage pressure on the GPU and host memory. During gradient accumulation operations, according to the preset number of gradient accumulation steps, after each forward and backward propagation calculation of a small batch is completed, to reduce memory pressure, the obtained individual gradient data is written into the expansion card in FP16 format.
[0028] Before performing the current gradient accumulation operation, it is necessary to read the previously saved FP16 gradient data from the expansion card. Among them, the previously saved FP16 gradient data refers to the gradient data calculated and written into the expansion card individually each time before the preset number of gradient accumulation steps (i.e., before gradient accumulation operations) is reached. These gradient data are intermediate results generated during the gradient accumulation process and are used for accumulation operations with newly calculated FP16 gradient data in subsequent steps.
[0029] Step S30: Accumulate the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data.
[0030] In this embodiment, the current FP16 gradient data is the gradient data obtained through the forward and backward propagation calculations of the current small batch of data and is stored in FP16 format.
[0031] To implement the gradient accumulation operation, read the previously saved FP16 gradient data from the expansion card and add the previously saved FP16 gradient data to the current FP16 gradient data. This operation can be achieved through the tensor operation interfaces provided by deep learning frameworks (such as PyTorch or TensorFlow). In deep learning frameworks, gradient data exists in the form of tensors, and the deep learning frameworks have built-in efficient tensor addition capabilities, enabling element-wise addition of tensors in FP16 format.
[0032] Exemplarily, in PyTorch, the torch.add() function is used to add two tensors in FP16 format, and the accumulated gradient data will continue to be stored in FP16 format.
[0033] The accumulated gradient data represents the sum obtained by accumulating the gradient data corresponding to all completed small batches of data in the current gradient accumulation step.
[0034] Step S40: Rewrite the accumulated gradient data back into the expansion card.
[0035] The gradient accumulation operation needs to calculate and accumulate the gradient data of multiple mini-batch data in the gradient accumulation steps. The accumulated gradient data needs to be saved for updating the optimizer parameter status. Since the GPU memory capacity is limited and all intermediate results cannot be stored in the GPU memory, the accumulated gradient data needs to be offloaded to the expansion card to relieve the storage pressure on the GPU and host memory.
[0036] In this embodiment, to write the accumulated gradient data back to the expansion card, the storage interface provided by the deep learning framework or hardware acceleration library can be called. For example, in PyTorch, gradient offloading can be achieved by moving the tensor data to the CPU or expansion card storage device.
[0037] After writing the accumulated gradient data back to the expansion card, in the subsequent steps of updating the optimizer parameter status, the accumulated gradient data can be read from the expansion card according to the update needs for further parameter update operations. After the update is completed, the separately calculated gradient data and the accumulated gradient data saved in the expansion card are cleared, and a new gradient accumulation cycle begins. Through the above method, multi-step gradient accumulation can be completed under limited memory conditions, thus supporting larger batch training.
[0038] In this embodiment, after the gradient accumulation operation is completed, the accumulated gradient data is rewritten back to the expansion card as the latest FP16 gradient data, and the latest FP16 gradient data will overwrite the previously saved FP16 gradient data and the current FP16 gradient data in the expansion card. It can be understood that when the gradient accumulation operation is completed, there is only one accumulated gradient data in the expansion card, and the accumulated gradient data is FP16 gradient data. Before the gradient accumulation operation is not executed, the separately calculated FP16 gradient data is saved in the expansion card.
[0039] Exemplarily, when the gradient accumulation step is set to 3, when calculating the first mini-batch to obtain the first gradient data, the first gradient data is written into the expansion card. Then calculate the second mini-batch to obtain the second gradient data, and continue to write the second gradient data into the expansion card. At this time, the preset gradient accumulation step has not been reached, and the first gradient data and the second gradient data are saved in the expansion card. When calculating the third gradient data of the third mini-batch, the third gradient data is also written into the expansion card, but it is determined that the preset gradient accumulation step has been reached through the judgment condition, so the gradient accumulation operation starts. Read the first gradient data and the second gradient data from the expansion card, then accumulate the first gradient data and the second gradient data with the third gradient data to obtain the accumulated gradient data. Finally, write the accumulated gradient data back to the expansion card and use the accumulated gradient data to update the optimizer status.
[0040] Exemplarily, to facilitate understanding of the implementation process of the gradient offloading optimization method based on DeepSpeed in this embodiment, please refer to Figure 2 , Figure 2 A brief schematic diagram of the gradient data flow based on DeepSpeed is provided. Specifically: During the deep learning model training phase, the storage locations of gradient data involve GPUs, CPUs (host memory), and expansion cards. During the backpropagation phase, FP16 gradient data for the current mini-batch of data is calculated on the GPU. After the FP16 gradient data is calculated, it is immediately offloaded to the expansion card instead of being retained in the host memory for gradient accumulation.
[0041] During the gradient accumulation phase, when a gradient accumulation operation needs to be performed, the previously saved FP16 gradient data is read from the expansion card and transferred to the GPU through the first buffer. Then, on the GPU, the currently calculated current FP16 gradient data is accumulated with the previously saved FP16 gradient data read from the expansion card. After the accumulation is completed, the accumulated gradient data is written back to the expansion card as the latest FP16 gradient data.
[0042] It should be noted that the previously saved FP16 gradient data refers to the gradient data that is individually calculated and written to the expansion card each time before the preset gradient accumulation step count is reached, that is, before the gradient accumulation operation is performed. The previously saved FP16 gradient data includes at least one FP16 gradient data, and its specific quantity is related to the gradient accumulation step count. For example, if the gradient accumulation step count is 3, the previously saved FP16 gradient data represents the FP16 gradient data calculated separately in the previous two times.
[0043] During the optimizer state update phase, the latest FP16 gradient data is read from the expansion card and converted to FP32 gradient data, and the FP32 gradient data is used to update the optimizer state in the host memory.
[0044] Based on the first embodiment of this application, in the second embodiment of this application, for the same or similar content as in the above-mentioned first embodiment, reference can be made to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , steps S101 to S103 are included before step S10: Step S101: Allocate a first buffer for the FP16 gradient data in advance.
[0045] During the training process of a deep learning model, the storage and management of gradient data significantly occupy memory resources. To efficiently process gradient data and avoid the performance overhead caused by frequent memory allocation and deallocation operations, a dedicated storage space is pre-allocated for FP16 gradient data.
[0046] The first buffer refers to a memory area with the data type of FP16 for storing gradient data in FP16 format. FP16 (half-precision floating-point number) is a data format that occupies less memory and is used in mixed-precision training to reduce memory occupancy and improve computational efficiency.
[0047] In this embodiment, the memory management interface provided by the deep learning framework or the hardware acceleration library can be called. For example, in PyTorch, a tensor of a specified size can be pre-allocated through the torch.empty() function or the torch.zeros() function, and the memory space occupied by this tensor represents the first buffer. In this way, the storage space can be prepared before the model training starts, avoiding dynamic memory allocation during the gradient calculation process, and thus significantly reducing the memory management overhead.
[0048] Step S102: Obtain the current FP16 gradient data and copy the current FP16 gradient data to the first buffer.
[0049] After the allocation of the first buffer is completed, it is necessary to obtain the current FP16 gradient data of the current mini-batch data and copy it to the first buffer. The current FP16 gradient data is the gradient data calculated through the forward propagation and backward propagation of the current mini-batch data and is stored in the FP16 data format.
[0050] In actual operation, through the tensor operation interface provided by the deep learning framework, the calculated current FP16 gradient data of the current mini-batch data can be copied to the first buffer to ensure the rapid storage of gradient data and provide a basis for subsequent gradient data transmission and gradient accumulation operations.
[0051] In this embodiment, by storing the gradient data in the pre-allocated first buffer, the performance loss caused by frequent memory allocation can be reduced, and at the same time, the overall efficiency of data processing can be improved.
[0052] Step S103: Write the current FP16 gradient data into the expansion card through the first buffer.
[0053] After copying the current FP16 gradient data to the first buffer, it is necessary to transfer the current FP16 gradient data from the first buffer to the expansion card for storage. As additional storage space, the expansion card can effectively relieve the pressure on the GPU and host memory. Especially in large-scale model training, the use of the expansion card can significantly reduce memory occupancy.
[0054] It should be noted that regardless of whether gradient accumulation is performed or not, the FP16 gradient data generated or calculated from each small batch of data will be written into the expansion card.
[0055] It should be noted that when performing gradient accumulation, the previously saved gradient data is read from the expansion card, and the previously saved gradient data is accumulated with the current FP16 gradient data calculated from the current small batch. This process does not directly accumulate all the gradient data saved in the expansion card with the current gradient data, but only accumulates up to the previously saved gradient data with the current gradient data. This can avoid repeatedly accumulating the same gradient data and ensure the accuracy and efficiency of the gradient accumulation operation.
[0056] Exemplarily, when the gradient accumulation step is set to 3, when calculating the first small batch to obtain the first gradient data, the first gradient data is written into the expansion card. Then, when calculating the second small batch to obtain the second gradient data, the second gradient data is continuously written into the expansion card. At this time, the preset gradient accumulation step has not been reached, and the first gradient data and the second gradient data are saved in the expansion card. When calculating the third gradient data of the third small batch, the third gradient data will still be written into the expansion card. At this time, it is determined that the preset gradient accumulation step has been reached through the judgment condition, so the gradient accumulation operation is started: the previously saved gradient data is read from the expansion card, that is, only the first gradient data and the second gradient data are read, and the third gradient data saved this time will not be read during this process. Then, the first gradient data, the second gradient data, and the third gradient data are accumulated to obtain the accumulated gradient data, and the accumulated gradient data is written into the expansion card.
[0057] In this embodiment, the current FP16 gradient data is written into the expansion card through the first buffer, and this process can be implemented through the data transfer interface provided by the deep learning framework or the hardware acceleration library. For example, in PyTorch, the torch.Tensor.to() method can be used to transfer the current FP16 gradient data from the host memory to the expansion card. In this way, not only can the gradient data be safely stored in the expansion card, but also the gradient data can be quickly read when needed later, thus realizing efficient gradient data management.
[0058] Exemplarily, to facilitate understanding of the implementation process of the gradient offloading optimization method based on DeepSpeed obtained by combining this embodiment with the above-mentioned first embodiment, when a gradient accumulation operation needs to be performed, a first buffer with a data type of FP16 data format is applied for. Subsequently, the gradient data saved last time is read from the expansion card and written into the first buffer. Then, in order to perform the gradient accumulation operation, the gradient data saved last time in the first buffer is transferred from the host memory to the GPU. At this time, the gradient data saved last time in the GPU is accumulated with the current FP16 gradient data calculated currently. In this way, the memory and video memory resources can be utilized efficiently, and the performance overhead caused by frequent memory allocation and release operations can be reduced.
[0059] In an optional implementation, after step S102, steps S1021 to S1022 may further be included: Step S1021: Determine the partition number, gradient offset, and gradient length of the current FP16 gradient data.
[0060] In this embodiment, in order to facilitate the management of the gradient data generated in the backpropagation stage, the gradient data is stored in partitions. The partition number indicates the partition to which the current FP16 gradient data belongs. Different partitions can store different types or different stages of gradient data, which is convenient for subsequent classification management and searching of the gradient data. The gradient offset indicates the starting position of the current FP16 gradient data within its own partition, determining the specific storage position of the current FP16 gradient data in the partition. The gradient length represents the storage space size occupied by the current FP16 gradient data, that is, the length of the current FP16 gradient data itself.
[0061] By determining the partition number, gradient offset, and gradient length of the current FP16 gradient data, the position and range of the current FP16 gradient data in the memory can be accurately located and described. In actual operation, the partition number, gradient offset, and gradient length and other information of the gradient data can be determined according to factors such as the source of the gradient data, the structure of the model, and the storage strategy. For example, the gradient data can be divided into different partitions according to different layers of the model, each partition corresponds to a partition number, the gradient offset is determined according to the order of the gradient data in this layer, and the gradient length is calculated according to the dimension and data type of the gradient data.
[0062] Step S1022: According to the partition number, the gradient offset, and the gradient length, write the current FP16 gradient data into the gradient file with the corresponding number through the first buffer, and the gradient file is located in the expansion card.
[0063] According to the partition number, gradient offset, and gradient length of the current FP16 gradient data, write the current gradient data into the gradient file with the corresponding number in the expansion card through the first buffer. The first buffer serves as a transfer area for gradient data and stores the current FP16 gradient data. Locate the target gradient file according to the partition number, then determine the starting position for writing the current FP16 gradient data according to the gradient offset, and finally write the current FP16 gradient completely into the gradient file according to the gradient length.
[0064] It should be noted that gradient files are all stored in the form of files in the expansion card, and the naming form is partition id + offset + gradient length.
[0065] In an optional implementation, the generated FP16 gradient data is written into the expansion card through the swap_out_gradients(parameter, gradient_offsets, gradient_tensors) method. This method contains three parameters: parameter is the parameter data corresponding to the current gradient data. In a deep learning model, each gradient data is associated with a certain parameter of the model. parameter stores the specific information of these parameters, provides context for the gradient data, and helps to determine the position and role of the gradient data in the model.
[0066] gradient_offsets is the offset of the current gradient data, representing the starting position of the current gradient data in the storage area of the expansion card. Specifying the offset can ensure that the current gradient data is accurately written to the corresponding position in the expansion card, avoiding data overwrite or confusion.
[0067] gradient_tensors is the current gradient data to be written into the expansion card.
[0068] In this embodiment, when calling the swap_out_gradients method, write the current gradient_tensors to the corresponding storage position in the expansion card according to the specified offset, so as to achieve efficient storage and management of gradient data.
[0069] Based on the above embodiments of the present application, in the third embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 4 , step S30 includes steps S310~S320: Step S310: Accumulate the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data.
[0070] Through the tensor operation interfaces provided by deep learning frameworks (such as PyTorch or TensorFlow), for example, using the torch.add() function, the previously saved FP16 gradient data read from the expansion card is element-wise added to the currently calculated FP16 gradient data. The accumulated gradient data contains the sum of the previously saved gradient data and the currently calculated current FP16 gradient data, providing complete gradient information for subsequent model parameter updates.
[0071] In an optional implementation, before performing the gradient accumulation operation, a temporary buffer in FP16 data format is first applied for temporarily storing the gradient data during the accumulation operation. This temporary buffer is a buffer in FP16 data format pre-allocated for the optimizer state update phase. When the optimizer state update operation has not been performed yet, this temporary buffer is applied for and temporarily used to perform the gradient accumulation operation. Specifically, when performing the gradient accumulation operation, the previously saved gradient data is read from the expansion card and copied to the temporary buffer. At the same time, the currently calculated current FP16 gradient data is also transferred to the temporary buffer for the gradient accumulation operation. In this way, the gradient accumulation operation can be efficiently processed before the optimizer state update, avoiding additional memory allocation.
[0072] Step S320: Rewrite the accumulated gradient data back to the corresponding gradient file in the expansion card according to the partition number, gradient offset, and gradient length.
[0073] In this embodiment, after the gradient accumulation operation is completed, a write request is generated through the file operation interface provided by the deep learning framework or the operating system. Based on this write request, the accumulated gradient data is written into the corresponding named gradient file in the expansion card according to the partition number, gradient offset, and gradient length to implement the write-back operation. For example, the seek() method of the file is used to locate the position specified by the gradient offset, and then the write() method is used to write the accumulated gradient data into the corresponding gradient file. This method ensures the orderly storage of the gradient data in the expansion card. At the same time, when the gradient data needs to be used subsequently, the gradient file can be quickly and accurately read according to the partition number, gradient offset, and gradient length, improving the data read and write efficiency and making the gradient data easier to manage and maintain. At the same time, writing the accumulated gradient data back to the expansion card releases the space of the host memory and GPU video memory, providing more memory resources for the subsequent model training process and ensuring the efficient progress of the model training.
[0074] In this embodiment, after the gradient accumulation operation is completed, the accumulated gradient data is rewritten back to the expansion card as the latest FP16 gradient data. The latest FP16 gradient data will overwrite the original gradient data in the expansion card. The original gradient data refers to the previously saved FP16 gradient data and the current FP16 gradient data. This means that after the gradient accumulation operation is completed, only one accumulated gradient data is saved in the expansion card, and the accumulated gradient data is FP16 gradient data. Before the gradient accumulation operation is performed, the FP16 gradient data calculated separately each time is saved in the expansion card.
[0075] Exemplarily, to facilitate understanding of the implementation process of the gradient offloading optimization method based on DeepSpeed in this embodiment, please refer to Figure 5 , Figure 5 which provides a schematic diagram of the brief process in the backpropagation stage. Specifically: In the backpropagation stage, first, the FP16 gradient data of the current mini-batch of data is calculated, and it is determined whether the gradient accumulation operation needs to be performed. If the gradient accumulation operation is performed, the previously saved FP16 gradient data is read from the expansion card and copied to the temporary buffer. The temporary buffer is used to temporarily store the gradient data during the accumulation operation for accumulation with the currently calculated FP16 gradient data. The accumulated gradient data is rewritten back to the expansion card through the first buffer and overwrites the original gradient data in the expansion card. The original gradient data includes the previously saved gradient data and the current FP16 gradient data. After the gradient accumulation operation is completed, the optimizer state update step is completed by reading the latest gradient data in the expansion card, that is, the accumulated gradient data.
[0076] Based on the above embodiments of the present application, in the fourth embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, after step S40, there is also step S41: Step S41: When performing the optimizer state update, read the FP16 gradient data from the expansion card and convert the FP16 gradient data into FP32 gradient data to complete the optimizer state update.
[0077] During the training process of the deep learning model, the purpose of the optimizer state update is to adjust the parameters of the model according to the gradient data after gradient accumulation, so that the model can continuously learn and optimize to better fit the training data.
[0078] In this embodiment, the gradient data is stored in the expansion card in the FP16 data format, which can effectively relieve the pressure on the GPU and the host memory. Especially when dealing with large-scale model training and large datasets, it can avoid training interruption or performance degradation caused by insufficient memory.
[0079] When an optimizer state update is required, directly read the latest FP16 gradient data from the expansion card. The latest FP16 gradient data represents the result data after performing the gradient accumulation operation.
[0080] Exemplarily, by calling the swap_in_fp16_gradients() method, read the FP16 gradient data (i.e., the accumulated gradient data or the latest gradient data) finally saved to the expansion card during the backpropagation stage. In this method, the allocate_fp16() method will be called to allocate a memory space, i.e., the second buffer, for the FP16 gradient data to be read. The number of parameters that the second buffer can contain is the number of parameters in one partition.
[0081] In an optional implementation, steps S401 to S403 may also be included before step S41: Step S401: Allocate a second buffer and a third buffer. The second buffer is used to store the latest FP16 gradient data read from the expansion card, and the third buffer is used to store the converted FP32 gradient data.
[0082] To efficiently manage and convert gradient data, a second buffer and a third buffer are allocated for FP16 gradient data and FP32 gradient data respectively. The second buffer is used to temporarily store the latest FP16 gradient data read from the expansion card. The third buffer is used to store the gradient data obtained by converting the latest FP16 gradient data into the FP32 format. This buffer allocation strategy can ensure the efficient processing of gradient data during conversion and transmission, while reducing memory fragmentation.
[0083] During the training of a deep learning model, to better manage gradient data, the parameters of the model are usually divided into different partitions according to a preset rule, and the number of parameters contained in each partition is the number of partition parameters.
[0084] It should be noted that the number of partition parameters is not fixed and will be adjusted according to factors such as the model architecture, the characteristics of the training data, and the adopted partitioning strategy.
[0085] In this embodiment, according to the number of partition parameters, determine the sizes of the second buffer and the third buffer. Among them, the size of the second buffer is the number of partition parameters multiplied by 2 bytes, and the size of the third buffer is the number of partition parameters multiplied by 4 bytes. Through this calculation method, it can be ensured that the second buffer and the third buffer have sufficient space to store the FP16 gradient data and FP32 gradient data corresponding to the number of partition parameters.
[0086] Exemplarily, memory space is pre-allocated during the initialization phase of the SwapBufferManager class. A buffer of FP16 data type and at least four buffers of FP32 data type are allocated. By redefining the allocation and release functions, such as the allocate_all() function and the free() function, the buffers of the two data types, FP16 and FP32, are managed by judging the required data type.
[0087] Step S402: Read the latest FP16 gradient data from the expansion card, transfer the latest FP16 gradient data to the host memory through the second buffer, and convert the latest FP16 gradient data to the FP32 gradient data in the host memory.
[0088] Before performing the optimizer state update, read the latest FP16 gradient data from the expansion card and store the latest FP16 gradient data in the second buffer. Subsequently, the latest FP16 gradient data is transferred to the host memory and format conversion is performed in the host memory to convert the FP16 gradient data to FP32 gradient data. This process takes advantage of the mixed-precision training of DeepSpeed. By completing the conversion in the host memory, the performance loss caused by directly performing precision conversion on the GPU is avoided.
[0089] In the host memory, use the tensor data type conversion function provided by the deep learning framework. For example, use the to() method in PyTorch to convert the gradient data in FP16 data format to FP32 data format. The host memory has a relatively flexible and powerful computing environment and can efficiently execute the data type conversion task.
[0090] Exemplarily, perform type conversion through the float() method to convert the gradient data in FP16 data format to FP32 data format, and save the FP32 gradient data in the third buffer.
[0091] Step S403: Store the FP32 gradient data in the third buffer to complete the parameter update.
[0092] After completing the conversion of FP16 gradient data to FP32 gradient data, store the converted FP32 gradient data in the pre-allocated third buffer so that when the optimizer performs state update, the required FP32 gradient data can be quickly and accurately obtained from the third buffer.
[0093] When the optimizer performs a status update operation, it adjusts the model parameters based on the FP32 gradient data and its own update algorithm. For example, common optimizers such as Stochastic Gradient Descent (SGD), Adagrad (Adaptive Gradient), and Adam optimization algorithm will use the FP32 gradient data to calculate the update amount of the parameters, and then update the model parameters to complete the training and optimization process of the model.
[0094] Exemplarily, to facilitate understanding of the implementation process of the gradient offloading optimization method based on DeepSpeed in this embodiment, please refer to Figure 6 , Figure 6 which provides a schematic diagram of the brief process of the optimizer status update stage. Specifically: After the gradient accumulation operation is completed, the latest gradient data, that is, the accumulated gradient data, is currently stored in the expansion card. The latest gradient data in the expansion card is read into a pre-allocated second buffer, and the latest gradient data is converted into FP32 gradient data in the host memory, and the converted FP32 gradient data is saved to the third buffer for optimizer status update.
[0095] The embodiment of the present application provides a gradient offloading optimization device based on DeepSpeed. The gradient offloading optimization device based on DeepSpeed includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the gradient offloading optimization method based on DeepSpeed in the first embodiment above.
[0096] Next, refer to Figure 7 which shows a schematic structural diagram of a gradient offloading optimization device based on DeepSpeed suitable for implementing the embodiment of the present application. The gradient offloading optimization device based on DeepSpeed in the embodiment of the present application may include various hardware and software components for implementing the gradient offloading optimization method based on DeepSpeed. Figure 7 The shown gradient offloading optimization device based on DeepSpeed is only an example and should not bring any limitations to the functions and usage scope of the embodiment of the present application.
[0097] As Figure 7As shown, the DeepSpeed-based gradient offloading optimization device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the DeepSpeed-based gradient offloading optimization device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the DeepSpeed-based gradient offloading optimization device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a DeepSpeed-based gradient offloading optimization device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0098] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0099] The gradient offloading optimization device based on DeepSpeed provided by this application adopts the gradient offloading optimization method based on DeepSpeed in the above-mentioned embodiment, and can solve the technical problem of excessive video memory occupation during the training of large-scale deep learning models by DeepSpeed. Compared with the prior art, the beneficial effects of the gradient offloading optimization device based on DeepSpeed provided by this application are the same as those of the gradient offloading optimization method based on DeepSpeed provided by the above-mentioned embodiment, and other technical features in the gradient offloading optimization device based on DeepSpeed are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0100] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0101] As mentioned above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0102] The embodiment of this application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the gradient offloading optimization method based on DeepSpeed in the above-mentioned embodiment.
[0103] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0104] The above computer-readable storage medium can be included in the gradient offloading optimization device based on DeepSpeed; or it can exist independently without being assembled into the gradient offloading optimization device based on DeepSpeed.
[0105] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the gradient offloading optimization device based on DeepSpeed, the gradient offloading optimization device based on DeepSpeed is caused to: in response to performing a gradient accumulation operation, obtain current FP16 gradient data; read the previously saved FP16 gradient data from the expansion card; accumulate the previously saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data; and rewrite the accumulated gradient data back to the expansion card.
[0106] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN, Local Area Network) or a wide area network (WAN, Wide Area Network), or it can be connected to an external computer (for example, by connecting through the Internet service provider via the Internet).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0108] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0109] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned DeepSpeed-based gradient offloading optimization method, which can solve the technical problem of excessive video memory occupancy during large-scale deep learning model training using DeepSpeed. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the DeepSpeed-based gradient offloading optimization method provided in the above embodiments, and will not be elaborated here.
[0110] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned DeepSpeed-based gradient offloading optimization method.
[0111] The computer program product provided by the present application can solve the technical problem of excessive video memory occupation during the training of large-scale deep learning models by DeepSpeed. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as those of the DeepSpeed-based gradient offloading optimization method provided by the above embodiment, and will not be elaborated here.
[0112] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent scope of the present application.
[0113] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0115] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. A gradient unloading optimization method based on DeepSpeed, characterized in that: The DeepSpeed-based gradient unloading optimization method includes: In response to performing a gradient accumulation operation, obtaining current FP16 gradient data; Read the last saved FP16 gradient data from the expansion card; Accumulate the last saved FP16 gradient data with the current FP16 gradient data to obtain the accumulated gradient data; The accumulated gradient data is rewritten back into the expansion card.
2. The DeepSpeed-based gradient unloading optimization method according to claim 1, characterized in that: Before the step of obtaining the current FP16 gradient data in response to performing the gradient accumulation operation, the method further comprises: Pre-allocate the first buffer for FP16 gradient data; Obtain current FP16 gradient data, and copy the current FP16 gradient data to the first buffer; The current FP16 gradient data is written into the expansion card through the first buffer.
3. The DeepSpeed-based gradient unloading optimization method according to claim 2, characterized in that: After the step of obtaining the current FP16 gradient data and copying the current FP16 gradient data to the first buffer, the method further includes: Determine the partition number, gradient offset, and gradient length of the current FP16 gradient data; According to the partition number, the gradient offset, and the gradient length, the current FP16 gradient data is written into a gradient file with a corresponding number through the first buffer, and the gradient file is located in the expansion card.
4. The DeepSpeed-based gradient unloading optimization method according to claim 1, characterized in that: The step of accumulating the last saved FP16 gradient data and the current FP16 gradient data to obtain the accumulated gradient data includes: Accumulate the last saved FP16 gradient data with the current FP16 gradient data to obtain the accumulated gradient data; The accumulated gradient data is rewritten back to the corresponding gradient file in the expansion card according to the partition number, gradient offset and gradient length.
5. The DeepSpeed-based gradient unloading optimization method according to claim 1, characterized in that: The step of rewriting the accumulated gradient data back into the expansion card comprises: After the gradient accumulation operation is completed, the accumulated gradient data is written into the expansion card as the latest FP16 gradient data, and the latest FP16 gradient data will overwrite the last saved FP16 gradient data and the current FP16 gradient data in the expansion card.
6. The DeepSpeed-based gradient unloading optimization method according to claim 1, characterized in that: After the step of rewriting the accumulated gradient data back into the expansion card, the method further includes: When performing optimizer parameter update, the latest FP16 gradient data is read from the expansion card, and the latest FP16 gradient data is converted into FP32 gradient data to complete the optimizer parameter update.
7. The DeepSpeed-based gradient unloading optimization method according to claim 6, characterized in that: Before the step of reading the latest FP16 gradient data from the expansion card and converting the latest FP16 gradient data into FP32 gradient data to complete the step of updating the optimizer parameters when performing the optimizer parameter update, the method further includes: Allocating a second buffer and a third buffer, wherein the second buffer is used to store the latest FP16 gradient data read from the expansion card, and the third buffer is used to store the converted FP32 gradient data; Reading the latest FP16 gradient data from the expansion card, transmitting the latest FP16 gradient data to the host memory through the second buffer, and converting the latest FP16 gradient data into the FP32 gradient data in the host memory; The FP32 gradient data is stored in the third buffer to complete parameter update.
8. The DeepSpeed-based gradient unloading optimization method according to claim 7, characterized in that: The step of allocating a second buffer and a third buffer for the FP16 gradient data and the FP32 gradient data, respectively, further includes: The sizes of the second buffer and the third buffer are determined according to the number of partition parameters, wherein the size of the second buffer is the number of partition parameters multiplied by 2 bytes, and the size of the third buffer is the number of partition parameters multiplied by 4 bytes.
9. A gradient unloading optimization device based on DeepSpeed, characterized in that: The DeepSpeed-based gradient offloading optimization device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the DeepSpeed-based gradient offloading optimization method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the DeepSpeed-based gradient unloading optimization method are implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Text generation method and device, equipment and storage medium
CN117875367A
Single-GPU large model training method and system based on multiple SSDs
CN118939434A
Gradient updating method and device, equipment, medium and program product
CN119168007A
Layered Gradient Accumulation and Modular Pipeline Parallelism for Improved Training of Machine Learning Models
US20220383084A1
Training method for neural network model, and related product
WO2021244354A1