A method and device for optimizing video memory for model training
By recording and analyzing the video memory usage of each tensor, combining the swap-in and out and recalculation strategies, the video memory usage in model training is optimized, which solves the problems of low memory utilization and slow training speed, and reduces the video memory usage and improves the training speed.
Patent Information
- Application Number
- CN202411715255.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The existing technology has low memory utilization and slow training speed in model training, and lacks effective solutions.
By training a round based on the preset network model, the input time, output time, memory usage and calculation time of each tensor are recorded, and the transfer time overhead of each tensor from GPU video memory to main memory is calculated. The large top heap is preferred to swap in and out, and the recalculation method is combined with the method of recalculation to further reduce the memory usage.
Effectively reduce video memory usage, improve model training speed, and ensure that large-scale network models can be trained on target devices.
Smart Images

Figure CN119205484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a method and device for optimizing video memory for model training. Background Art
[0002] Deep learning and AI are developing rapidly, especially in the fields of computer vision, natural language processing, and recommendation systems. As the tasks become more complex, the models are getting larger and larger, occupying more and more video memory, while the memory capacity of GPUs is limited, which limits the scale of the models and the speed of training.
[0003] There are three commonly used methods for video memory optimization: quantization (compression, mixed precision), recomputation, and swapping in and out. Quantization is the conversion of high-precision values in the model into low-precision values; recomputation is a strategy that does not store all intermediate activation values during forward propagation, but recalculates these values during backward propagation; swapping in and out is a way to manage video memory usage by dynamically transferring data between video memory and main memory. For example, the Chinese patent document with publication number CN112329834A discloses a method and device for allocating video memory space during recurrent network model training, which effectively compresses the video memory used in network calculations, thereby increasing the training speed; the Chinese patent document with publication number CN115437795A discloses a method and system for optimizing video memory recomputation based on load perception of heterogeneous GPU clusters, which determines the stage with the highest video memory load among all stages, performs recomputation optimization based on an algorithm that minimizes video memory overhead, and ensures load balance at each stage.
[0004] All of the above methods can reduce the usage of video memory, but there is a problem of reducing accuracy or training efficiency. There is currently no effective solution to the problem of low video memory utilization and slow training speed in model training in related technologies. Summary of the invention
[0005] The present invention provides a method and device for optimizing video memory for model training, which can effectively reduce video memory occupancy and improve the training speed of the model.
[0006] A method for optimizing video memory for model training, comprising:
[0007] (1) Based on the preset network model, train one round and record the input time, output time, video memory usage, and computing time of each tensor;
[0008] (2) Based on the transmission speed and memory usage, calculate the transfer time overhead of each tensor from GPU memory to main memory; based on the input time and output time of each tensor, obtain a list of hidden transfer overheads, which includes the tensor ID and memory usage;
[0009] (3) Based on the list of hidden transfer overheads, a large top heap is established based on the memory usage of the tensor, and the tensors at the top of the heap are selected in turn as candidate tensors for swapping in and out, thereby obtaining a list of candidates for swapping in and out;
[0010] (4) When swapping in and out while hiding the transfer overhead cannot meet the video memory demand, recalculation is used to further reduce the video memory usage and obtain a list of candidates for recalculation;
[0011] (5) Further optimization is performed based on the memory optimization strategy of steps (3) and (4) to enable large-scale network models to be trained on the target device.
[0012] In step (1), during training, the batch size of the network input is set to 1.
[0013] The specific process of step (2) is as follows:
[0014] Transfer the GPU memory to the main memory, and obtain the data transfer speed by recording the time overhead of the specified data transfer size;
[0015] According to the data transmission speed and the memory usage of the tensor, the transfer time overhead of each tensor from the GPU memory to the main memory is calculated;
[0016] According to the access time of each tensor, if the tensor output time plus 2 times the transfer time overhead is less than the next access time of the tensor, the tensor is judged to be a tensor that can hide the transfer overhead, thereby obtaining a list of tensor IDs and video memory occupancy.
[0017] In step (3), the tensors at the top of the stack are selected in turn as candidate tensors for swapping in and out, and a list of candidates for swapping in and out is obtained, specifically:
[0018] The amount of memory that needs to be reduced is obtained by subtracting the GPU memory required by the network before optimization and the currently available memory. Then, the loop is traversed to take out the element from the top of the heap and mark the tensor as selected for swapping in and out. After the traversal is completed, a list of candidates for swapping in and out is obtained, which contains the tensor ID and the memory usage.
[0019] The end condition of the loop traversal is that the sum of the memory usage of the swapped-in and swapped-out tensors is greater than or equal to the required memory usage reduction or the heap is empty;
[0020] If the loop ends and the sum of the memory usage of the swapped-in and swapped-out tensors is greater than or equal to the required memory usage reduction, the memory optimization process ends and the process goes to step (5). Otherwise, the process goes to step (4) and further optimizes the memory by combining recalculation means.
[0021] In step (4), the recalculation method is combined to further reduce the memory usage, and a list of candidates for recalculation is obtained, specifically:
[0022] Traverse the computation graph, and for each operator, determine whether the output tensor of its predecessor operator is selected for swapping in and out. If it is selected, the memory usage of the operator minus the memory usage reduced by the predecessor operator through swapping in and out, and then divided by the operator calculation time to get an evaluation index value; if it is not selected, directly divide the operator memory usage by the operator calculation time to get the value; in this way, get the corresponding list, which includes the operator id and the evaluation index value;
[0023] According to the value of the evaluation index, a large top heap is established;
[0024] Take them out from the top of the heap one by one and mark them as recalculated until the reduced video memory usage is less than or equal to the currently available video memory;
[0025] After the traversal is completed, a list of candidates for recalculation is obtained, which contains the operator ID and the memory usage.
[0026] In step (5), further optimization processing is performed according to the video memory optimization strategy of steps (3) and (4), specifically:
[0027] For each tensor ID, determine whether it is in the candidate list and choose to implement the corresponding strategy;
[0028] If it is marked as recalculated, the video memory is released, and when the tensor is accessed again, it is calculated by the previous output tensor;
[0029] If the tensor is marked as swapped in and out, it will be transferred to the main memory after the tensor is obtained and calculated. The prefetch time is obtained based on the next access time and the transfer time overhead, and it will be transferred back to the GPU video memory before the prefetch time.
[0030] A device for optimizing video memory for model training, comprising:
[0031] The pre-evaluation module is used to obtain the input time, output time, memory usage, calculation time, and time overhead of transferring each tensor from GPU memory to main memory in a round of training.
[0032] The strategy selection module is used to select candidate tensors to be swapped in and out and candidate operators to be recalculated, and to determine an optimization strategy that combines recalculation and swapping in and out;
[0033] The training module is used to combine the deep learning framework during actual training, and implement the corresponding optimization method on the corresponding tensor or operator according to the optimization strategy obtained by the strategy selection module, so that the large-scale network model can be trained on the target device.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The model training video memory optimization method provided in this embodiment is based on a preset network model, trains for one round, and records the input time, output time, video memory occupancy and calculation time of each tensor; according to the input time and output time of each tensor, a list of transfer overheads that can be hidden is obtained, and when the transfer synchronization overhead can be avoided, the tensors are swapped in and out for video memory optimization. When swapping in and out cannot meet the video memory demand while hiding the transfer overhead, a new evaluation algorithm is used to select an operator for the recalculation strategy in combination with the recalculation strategy; a strategy with lower strategy selection time overhead is used to obtain a combination of swapping in and out and recalculation, which effectively reduces the video memory occupancy to meet the needs of large model training while ensuring the training efficiency as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 The present invention is a flowchart of a method for optimizing video memory for model training according to an embodiment of the present invention.
[0037] Figure 2 The figure is a schematic diagram of a specific optimization strategy selection process in an embodiment of the present invention.
[0038] Figure 3 A schematic diagram of the structure of a device for optimizing video memory for model training according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.
[0040] First, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0041] (1) Computational graphs are graphical representations of the computational process. They are a “language” for describing equations. Since they are graphs, they have nodes (variables) and edges (operations (simple functions)). In the field of deep learning, neural network models can essentially be represented by a computational graph, and their training process can be divided into three parts: forward propagation, back propagation, and parameter update.
[0042] (2) Video memory, also called frame buffer, is used to store data processed by the graphics chip or rendering data to be extracted. Like the computer's memory, video memory is a component used to store graphics information to be processed.
[0043] (3) Tensor is a multidimensional array used to represent and store data. It is the basic representation of data in deep learning frameworks (such as TensorFlow, PyTorch, etc.).
[0044] like Figure 1 As shown, a model training video memory optimization method includes the following steps:
[0045] Step S101, based on the preset network model, obtain the access time, video memory occupancy, and calculation time of each tensor in a round of training.
[0046] In some embodiments, the preset network model can be any type of network model to be trained, such as a deep neural network model to be trained, a residual network model to be trained, or any large-scale neural network model to be trained, etc. The batch size of the network input is set to 1, and one round of training is performed to obtain the input and output time, video memory occupancy, and computing time of each tensor.
[0047] In some possible implementations, obtaining the input and output times, memory usage, and computing time of each tensor requires the use of methods provided by a specific deep learning framework. Taking PyTorch as an example, PyTorch provides hook functions. By registering forward and backward propagation hooks, access times are recorded in the hook functions. The memory usage is evaluated based on the size and data type of each tensor. The computing time of a single operator can be obtained by subtracting the times recorded before and after the calculation.
[0048] Step S102, according to the transmission speed and the memory usage, the time cost of transferring each tensor from the GPU memory to the main memory is calculated, and a list of (tensor ID, memory usage) that can hide the transfer cost is obtained.
[0049] In some embodiments, the transmission speed is calculated based on the transmission data size and time using the interface provided by the acceleration chip GPU; the time overhead of transferring each tensor from the GPU video memory to the main memory is obtained based on the transmission speed and the video memory occupancy; based on the transfer time overhead and the access time, the tensors that can be swapped out to the main memory and swapped into the GPU video memory before being accessed again are the tensors that can hide the transfer overhead, and a list of (tensor IDs, video memory occupancy) that can hide the transfer overhead is obtained.
[0050] In some possible implementations, different acceleration chips will provide different interfaces. Taking NVIDIA GPU as an example, it provides cudaMemcpy, which can transfer GPU video memory to main memory. By recording the time overhead of specifying the data transmission size, the data transmission speed can be obtained; then, according to the video memory occupancy of the tensor obtained in step S101, the time overhead of transferring each tensor from GPU video memory to main memory is calculated; then, according to the access time of each tensor obtained in step S101, if the time of tensor output plus twice the transfer time overhead (swap in and swap out) is less than the next access time of the tensor, it is a tensor that can hide the transfer overhead, thereby obtaining a list of (tensor id, video memory occupancy) that can hide the transfer overhead.
[0051] Step S103, establishing a max-heap according to the amount of video memory that can be reduced, and selecting tensors as swap-in and swap-out candidate tensors.
[0052] In some embodiments, the memory usage of the tensor is the amount of memory that can be reduced, and a max-heap is established based on this value; the memory is taken out from the top of the heap in sequence and marked as swapped in and out until the reduced memory usage is less than or equal to the currently available memory or the heap is empty; if training is possible at this time, the memory optimization process is terminated; if the memory is still insufficient, the memory is further optimized in combination with recalculation means.
[0053] In some possible implementations, according to the heap-related interface provided by the programming language used, the list of (tensor ids, video memory usage) that can hide the transfer overhead obtained in step S102 is used to establish a max-heap with the video memory usage; then the video memory usage that needs to be reduced is obtained by subtracting the GPU video memory required by the network before optimization and the currently available video memory, and then a loop is traversed to take out elements from the top of the heap and mark the tensor as selected for swapping in and out. The condition for the end of the loop is that the sum of the video memory usage of the swapped-in and swapped-out tensors is greater than or equal to the video memory usage that needs to be reduced or the heap is empty; if the loop ends and the sum of the video memory usage of the swapped-in and swapped-out tensors is greater than or equal to the video memory usage that needs to be reduced, then the process of selecting the video memory optimization strategy ends and the process goes to step S105; otherwise, the process goes to step S104 to further optimize the video memory by combining the recalculation method.
[0054] Step S104, when swapping in and out while hiding the overhead cannot meet the video memory demand, recalculation means are combined to further reduce the video memory occupancy.
[0055] In some embodiments, when swapping in and out cannot meet the video memory demand while hiding the overhead, the computation graph is traversed, and for each operator, it is determined whether the output tensor of its predecessor operator is selected for swapping in and out. If selected, the video memory occupancy of the operator is subtracted from the video memory occupancy reduced by the predecessor operator through swapping in and out, and then divided by the operator calculation time to obtain an evaluation index value. If not selected, the video memory occupancy of the operator is directly divided by the operator calculation time to obtain the value, thereby obtaining the corresponding (operator id, evaluation index value) list; according to the size of the evaluation index value, a max-heap is established; and the operators are taken out from the top of the heap in turn and marked as recalculated until the reduced video memory occupancy is less than or equal to the currently available video memory.
[0056] In some possible implementations, after step S103, there is a list of (tensor ids, video memory usage) that are candidates for swapping in and out, and the evaluation value of each operator is calculated according to the proposed evaluation value calculation formula and a max-heap is established: determine whether the output tensor of its predecessor operator is selected for swapping in and out, if selected, the video memory usage of the operator minus the video memory usage reduced by the predecessor operator through swapping in and out, and then divided by the operator calculation time; if not selected, directly divide the operator video memory usage by the operator calculation time; when selecting a candidate operator for recalculation in the max-heap, if the output tensor of the predecessor operator of the operator is a candidate for swapping in and out, the tensor needs to be removed from the swapping in and out candidate list; after the traversal is completed, a list of (operator ids, video memory usage) that are candidates for recalculation is obtained;
[0057] Step S105, performing optimization processing according to the obtained video memory optimization strategy, so that a large-scale network model can be trained on the device.
[0058] In some possible implementations, after steps S103 and S104, there is a list 1 of (tensor id, video memory occupancy) as candidates for swapping in and out and a list 2 of (operator id, video memory occupancy) as candidates for recalculation. If, after the swapping in and out strategy of step S103, the optimized video memory is smaller than the currently available video memory, list 2 is empty. During the video memory optimization process, during the network training process, for each tensor id, it is determined whether it is in the candidate list, and a corresponding strategy is selected for implementation: if it is marked as recalculation, the video memory is released, and when the tensor is accessed again, it is calculated by the pre-placed output tensor; if it is marked as swapping in and out, after the tensor is obtained and calculated, it is transferred to the main memory, and the pre-fetching time is obtained according to the next access time and the transfer time overhead, and the pre-fetching time is re-transferred to the GPU video memory before the pre-fetching time.
[0059] like Figure 2 As shown, the specific optimization strategy selection of the model training video memory optimization method in the embodiment of the present invention may include the following steps:
[0060] Step S201 , according to a list of (tensor IDs, video memory usage) that can hide transfer overhead, a max-heap is established with video memory usage, and the heap is traversed in a loop.
[0061] In some embodiments, it is necessary to set an id and a flag for each tensor in the computation graph to facilitate determining the optimization strategy for each tensor.
[0062] Step S202, take out the top element of the heap, mark the tensor corresponding to id as a candidate for swapping in and out, and update the heap.
[0063] In some embodiments, according to the characteristics of the max-heap, the top element of the heap is the tensor with the largest memory usage, and the tensor corresponding to its ID is marked as swapped in and out; the operation of updating the heap depends on the specific computer language used, and the top element of the heap is swapped with the last element of the heap, the last element of the heap is taken out and the top element of the heap is sunk to a position that satisfies the characteristics of the heap to complete the update of the heap.
[0064] Step S203, entering different stages according to whether the total amount of video memory swapped in and out is greater than or equal to the amount of video memory that needs to be reduced.
[0065] In some embodiments, if the sum of the video memory swapped in and out is greater than or equal to the amount of video memory that needs to be reduced, the video memory optimization strategy selection phase is directly terminated and the process proceeds to step S213, that is, the final video memory optimization strategy is obtained.
[0066] In some embodiments, the total amount of video memory swapped in and out is less than the amount of video memory that needs to be reduced, and the process goes to step S204.
[0067] Step S204, determine whether the heap is empty. If the heap is not empty, return to step S202 and continue to select. If the heap is empty, enter step S205, combine the recalculation strategy to further reduce the memory usage, and traverse each operator in the calculation graph.
[0068] Step S206, performing different evaluation value calculation formulas according to whether the output tensor of the preceding operator of the operator is selected as swapped in or out.
[0069] In some embodiments, if the output tensor of its predecessor operator is selected for swapping in and out, the process goes to step S207, where the memory occupancy of the operator is subtracted from the memory occupancy of the predecessor operator reduced by swapping in and out, and then divided by the operator calculation time to obtain an evaluation index value, which is put into the (operator id, evaluation index value) list.
[0070] In some embodiments, if the output tensor of its predecessor operator is not selected for swapping in or out, the process proceeds to step S208, where the operator memory occupancy is divided by the operator calculation time to obtain an evaluation index value, which is put into the (operator id, evaluation index value) list.
[0071] Step S209 , according to the (operator id, evaluation index value) list, a max-heap is established with the evaluation index value, and the heap is traversed in a loop.
[0072] Step S210, take out the top element of the heap, mark the tensor corresponding to id, select it for recalculation, and update the heap.
[0073] In some embodiments, according to the characteristics of the max-max heap, the top element of the heap is the tensor with the largest memory occupancy, and the tensor corresponding to its ID is marked as recalculated; when recalculation is selected, if the output tensor of the preceding operator of the operator is a candidate for swapping in and out, the tensor needs to be removed from the swap-in and swap-out candidate list.
[0074] Step S211, entering different stages according to whether the total amount of video memory reduced by the swap-in, swap-out and recalculation strategy is greater than or equal to the amount of video memory that needs to be reduced.
[0075] In some embodiments, the total amount of video memory reduced by the swap-in, swap-out and recalculation strategies is greater than or equal to the amount of video memory required to be reduced, and the video memory optimization strategy selection phase ends, and the process proceeds to step S213, that is, obtaining the final video memory optimization strategy.
[0076] In some embodiments, the total amount of video memory reduced by the swap-in, swap-out and recalculation strategies is less than the amount of video memory that needs to be reduced, and the process goes to step S212.
[0077] Step S212, determine whether the heap is empty. If the heap is not empty, return to step S210 and continue to select. If it is empty, enter step S214 and output that no available video memory optimization strategy is found.
[0078] Step S213, obtaining the final video memory optimization strategy and entering the optimization processing stage.
[0079] Step S214: No available video memory optimization strategy is found, indicating that the network model cannot be trained on the target hardware through the video memory optimization method of swapping in and out combined with recalculation.
[0080] like Figure 3 As shown, a model training video memory optimization device includes: a pre-evaluation module 301, a strategy selection module 302 and a training module 303, wherein:
[0081] The pre-evaluation module 301 is used to obtain the access time, memory usage, calculation time of each tensor in a round of training, and the time overhead of transferring each tensor from the GPU memory to the main memory.
[0082] The strategy selection module 302 is used to select candidate tensors to be swapped in and out and candidate operators to be recalculated, and determine a comprehensive recalculation and swap-in and swap-out combination strategy.
[0083] The training module 303 is used to combine the deep learning framework during actual training, and implement the corresponding optimization method on the corresponding tensor or operator according to the optimization strategy obtained by the strategy selection module, so that the large-scale network model can be trained on the target device.
[0084] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for optimizing video memory for model training, characterized in that: include: (1) Based on the preset network model, train one round and record the input time, output time, video memory usage, and computing time of each tensor; (2) Based on the transmission speed and the memory usage, calculate the transfer time overhead of each tensor from the GPU memory to the main memory; based on the input time and output time of each tensor, if the tensor output time plus 2 times the transfer time overhead is less than the next access time of the tensor, then the tensor is judged to be a tensor that can hide the transfer overhead, thereby obtaining a list of the tensor that can hide the transfer overhead, which contains the tensor ID and the memory usage; (3) Based on the list of hidden transfer overheads, a large top heap is established based on the memory usage of the tensor, and the tensors at the top of the heap are selected in turn as candidate tensors for swapping in and out, thereby obtaining a list of candidates for swapping in and out; specifically: The amount of memory that needs to be reduced is obtained by subtracting the GPU memory required by the network before optimization and the currently available memory. Then, the loop is traversed to take out the element from the top of the heap and mark the tensor as selected for swapping in and out. After the traversal is completed, a list of candidates for swapping in and out is obtained, which contains the tensor ID and the memory usage. The end condition of the loop traversal is that the sum of the memory usage of the swapped-in and swapped-out tensors is greater than or equal to the required memory usage reduction or the heap is empty; if the loop ends and the sum of the memory usage of the swapped-in and swapped-out tensors is greater than or equal to the required memory usage reduction, the memory optimization process ends and the process goes to step (5); otherwise, the process goes to step (4) and further optimizes the memory by combining recalculation means; (4) When swapping in and out while hiding the transfer overhead cannot meet the video memory demand, recalculation is used to further reduce the video memory usage and obtain a list of candidates for recalculation; specifically: Traverse the computation graph, and for each operator, determine whether the output tensor of its predecessor operator is selected for swapping in and out. If it is selected, the memory usage of the operator minus the memory usage reduced by the predecessor operator through swapping in and out, and then divided by the operator calculation time to get an evaluation index value; if it is not selected, directly divide the operator memory usage by the operator calculation time to get the value; in this way, get the corresponding list, which includes the operator id and the evaluation index value; According to the value of the evaluation index, a large top heap is established; Take them out from the top of the heap one by one and mark them as recalculated until the reduced video memory usage is less than or equal to the currently available video memory; After the traversal is completed, a list of candidates for recalculation is obtained, which contains the operator ID and the amount of video memory occupied; (5) According to the video memory optimization strategy of steps (3) and (4), further optimization processing is performed to enable large-scale network models to be trained on the target device; specifically: For each tensor ID, determine whether it is in the candidate list and choose to implement the corresponding strategy; If it is marked as recalculated, the video memory is released, and when the tensor is accessed again, it is calculated by the previous output tensor; If the tensor is marked as swapped in and out, it will be transferred to the main memory after the tensor is obtained and calculated. The prefetch time is obtained based on the next access time and the transfer time overhead, and it will be transferred back to the GPU video memory before the prefetch time.
2. The model training video memory optimization method according to claim 1, characterized in that: In step (1), during training, the batch size of the network input is set to 1.
3. The model training video memory optimization method according to claim 1, characterized in that: In step (2), the GPU video memory is transferred to the main memory, and the data transmission speed is obtained by recording the time overhead of specifying the data transmission size; based on the data transmission speed and the video memory occupancy of the tensor, the transfer time overhead of each tensor from the GPU video memory to the main memory is calculated.
Citation Information
Patent Citations
Video memory space distribution method and device during loop network model training
CN112329834A
Video memory recalculation optimization method and system for load awareness of heterogeneous GPU (Graphics Processing Unit) cluster
CN115437795A
GPU memory optimization method and system for a deep learning training task
CN109919310A
Tensor-based deep learning GPU memory management optimization method and system
CN111078395A