A Memory Management Method and System for Deep Neural Networks
The method optimizes DNN memory management by prioritizing re-computation and memory swapping based on tensor dependencies and hardware parameters, reducing memory usage and training time, and minimizing fragmentation.
Patent Information
- Application Number
- CN202310161062.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-24
AI Technical Summary
During the existing deep neural network training process, the memory utilization rate is low, the training time is long, and memory swap causes memory fragmentation problems, limiting the optimization space for memory usage.
By calculating the tensor's eviction time and recalculation time ratio based on the tensor's data dependencies and hardware parameters before training, selecting appropriate recalculation or memory exchange methods, optimizing the storage state of the tensor, and using different storage strategies during the training process to avoid memory fragmentation.
It improves the memory utilization rate, reduces training time, solves the memory fragmentation problem, increases the sample size per iteration, and improves the overall training efficiency.
Smart Images

Figure CN116107754B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data storage, and more specifically, relates to a memory management method and system for deep neural networks. Background Technique
[0002] Deep neural networks (DNNs) are widely used in many machine learning tasks, such as computer vision, natural language processing, etc. To obtain higher training accuracy, the number of layers of DNN models used for training is getting deeper and the scale is getting larger. The existing GPU memory size cannot meet the storage overhead required for training common DNN models. Usually, training can only be completed by reducing the batch size, which not only increases the training time but also greatly reduces the utilization rate of the GPU.
[0003] During the training process of deep learning, the memory occupation mainly includes feature maps, gradient maps, convolution workspaces, and weights. Among them, the weights account for a small proportion, the gradient maps and convolution workspaces will be released immediately after the current calculation is completed, while the feature maps account for a large proportion and will reside in the video memory for a long time. Therefore, the existing work mainly optimizes the video memory occupation of feature maps (i.e., tensors).
[0004] The existing work mainly releases the memory of some tensors in the idle state during the forward propagation process through recomputation and memory swapping, and regenerates the tensor during the backward propagation process. However, the existing work selects the methods of recomputation and memory swapping based on the characteristics of different layers of the neural network. For layers of the same type, tensors of different sizes will be generated, which makes the above coarse-grained decision not the best decision, limiting the optimization space of memory usage. During the memory swapping process, since some tensors are transferred to the main memory for storage, a large amount of memory fragmentation will occur in the video memory, making the memory released through memory swapping not well utilized during the subsequent training process, reducing the utilization rate of the video memory and the overall training time of the deep neural network is slower. Summary of the Invention
[0005] Aiming at the above defects or improvement requirements of the existing technology, the present invention provides a memory management method and system for deep neural networks, aiming to reduce the video memory occupation of training deep neural networks, increase the number of samples processed in one iteration, improve the utilization rate of memory, and reduce the overall training time of deep neural networks when the video memory capacity is limited.
[0006] To achieve the above object, in the first aspect, the present invention provides a memory management method for deep neural networks, including:
[0007] If the video memory capacity required for training a deep neural network is greater than the provided video memory capacity, then:
[0008] Perform the following operations before training:
[0009] A1. Set the storage state of each tensor of the deep neural network to the state of writing to the video memory, and denote the set composed of all tensors of the deep neural network as the tensor set;
[0010] A2. Based on the data dependence relationship of each tensor in the tensor set, calculate the recomputation time of each tensor in the tensor set; obtain the tensor with the largest eviction time in the tensor set and the tensor with the largest ratio of its size to the recomputation time. If the generation cost of the former is less than the generation cost of the latter, set the storage state of the former to the memory swapping state and remove it from the tensor set; otherwise, set the storage state of the latter to the recomputation state and remove it from the tensor set;
[0011] A3. Update the data dependence relationship of each tensor in the tensor set, and repeat steps A2 - A3 for iteration until the video memory capacity required for training the deep neural network is less than or equal to the provided video memory capacity;
[0012] During training: For the tensor with the storage state of writing to the video memory, store it in the video memory during the forward propagation process; for the tensor with the storage state of memory swapping, transfer it to the main memory during the forward propagation process and prefetch it from the main memory to the video memory before accessing the tensor during the backward propagation; for the tensor with the storage state of recomputation, do not store the tensor during the forward propagation process, but recompute the tensor in the video memory according to the data dependence relationship before accessing the tensor during the backward propagation;
[0013] Among them, the eviction time of the tensor is the time interval from when the tensor is accessed during the forward propagation to when it is accessed during the backward propagation minus the transmission time of the tensor when training the deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network; the generation cost of the former is the ratio of the tensor size to the PCIe bandwidth; the generation cost of the latter is the recomputation time of the tensor.
[0014] Further preferably, the transmission time of the tensor is the sum of the time to unload the tensor to the main memory and the time to prefetch the tensor from the main memory when training the deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network.
[0015] Further preferably, the recomputation time of the tensor is obtained by adding the computation times of the tensors it depends on based on the data dependence relationship in the tensor set.
[0016] Further preferably, during training: Allocate two storage blocks in the video memory, denoted as the default storage block and the swap storage block respectively;
[0017] For tensors with the storage state being the state of writing to the video memory, store them in the default storage block during the forward propagation process; for tensors with the storage state being the memory swapping state, write them to the swapping storage block during the forward propagation process, and transfer them from the swapping storage block to the main memory when the PCIe resources are idle. Before accessing the tensors during the backward propagation, transfer them to the swapping storage block when the PCIe resources are idle; for tensors with the storage state being the recomputation state, do not store the tensors during the forward propagation process, but before accessing the tensors during the backward propagation, recompute the tensors in the video memory according to the data dependencies and store them in the default storage block.
[0018] Further preferably, the above step A2 includes:
[0019] Based on the data dependencies of the tensors in the tensor set, calculate the recomputation time of each tensor in the tensor set; calculate the ratio of the size of each tensor in the tensor set to its recomputation time to obtain the memory released per unit time corresponding to each tensor.
[0020] Sort the tensors in the tensor set in descending order of their eviction time to obtain the memory swapping priority sequence; sort the tensors in the tensor set in descending order of the memory released per unit time corresponding to them to obtain the recomputation priority sequence.
[0021] Obtain the first tensors in the memory swapping priority sequence and the recomputation priority sequence, and compare their generation overheads. If the former is less than the latter, set the storage state of the former to the memory swapping state and remove it from the tensor set; otherwise, set the storage state of the latter to the recomputation state and remove it from the tensor set.
[0022] Further preferably, the above memory management method further includes step A0 executed before step A1, which specifically includes: training the deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network, and recording the access time, calculation time, and size of each tensor.
[0023] Further preferably, the above deep neural network is a deep neural network for image processing or natural speech processing.
[0024] In a second aspect, the present invention provides a memory management system for a deep neural network, including: a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it executes the memory management method provided in the first aspect of the present invention.
[0025] In a third aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program. When the computer program is run by a processor, it controls the device where the storage medium is located to execute the memory management method provided in the first aspect of the present invention.
[0026] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0027] 1. The present invention provides a memory management method for deep neural networks. It estimates and compares the memory and computing resources consumed by the recomputation and memory swapping methods based on the computational graph of the deep neural network, the number of samples processed in one iteration, and the hardware parameters, and selects the method with lower overhead for different tensors. Since both the tensor size (memory occupied in the video memory) and the computing time (overall training time) are considered when sorting the overhead, but the sorting methods of the influencing factors of these two aspects for the recomputation and memory swapping methods are different and cannot be directly sorted and compared by a unified standard. Therefore, to achieve the above objectives, the present invention overall takes the training time as the priority optimization index, and selects the time for regenerating the tensor during backpropagation for the sorted tensors. Specifically, the present invention estimates the eviction time of the tensor and the memory released per unit time (the ratio of the tensor size to its recomputation time) based on the tensor size and hardware parameters, so as to obtain the two tensors that have the greatest impact on the video memory capacity under the memory swapping method and the recomputation method, that is, the tensor with the largest eviction time and the tensor with the largest memory released per unit time. By comparing the time overhead of regenerating (recomputing or prefetching) these two tensors, the method with lower overhead is selected to process the tensor. The present invention performs memory management based on tensor granularity, allowing different decisions to be triggered during tensor access. In the case of limited video memory capacity, it reduces the video memory occupancy for training deep neural networks, increases the number of samples processed in one iteration, improves the memory utilization rate, and greatly reduces the overall training time of the deep neural network.
[0028] 2. Further, during the training process, since memory swapping may cause memory fragmentation problems, the blocks released by the evicted tensors may not be able to be allocated and used in subsequent training processes. To further solve the technical problem of memory fragmentation caused by memory swapping, the memory management method provided by the present invention implements different placement strategies for tensors with a memory swapping storage state and tensors with a write-to-video-memory storage state. By placing them in different storage blocks respectively, it effectively avoids the memory fragmentation problem caused by memory swapping and further improves the memory utilization rate. Description of the Drawings
[0029] Figure 1Schematic diagram of the memory management method for deep neural networks provided in Embodiment 1 of the present invention;
[0030] Figure 2 Schematic diagram of the memory swapping overhead sorting method provided in Embodiment 1 of the present invention;
[0031] Figure 3 Schematic diagram of the fragmentation solution method provided in Embodiment 1 of the present invention. Detailed implementation manners
[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0033] To solve the problem of insufficient memory during the training of deep neural networks, recalculation and memory swapping methods are usually used. The specific practices include: the recalculation method is not to write tensors into the video memory during forward propagation, but to recalculate tensors according to data dependencies before accessing tensors during backward propagation; the memory swapping method is to transfer tensors to the main memory during forward propagation and prefetch them from the main memory to the video memory before accessing tensors during backward propagation.
[0034] Since these two methods will incur different overheads: the recalculation method recalculates tensors during backward propagation, requiring additional computing resources; the memory swapping method occupies the PCIe channel from the main memory to the video memory during backward propagation. Therefore, when the video memory capacity is insufficient, the present invention selectively selects the recalculation and memory swapping methods to process the intermediate data in each iteration during training; by comparing the overheads of the two methods, the method with lower overhead is selected to process tensors. The memory management method for deep neural networks provided by the present invention will be described in detail below with reference to specific embodiments:
[0035] Embodiment 1,
[0036] A memory management method for deep neural networks provided in this embodiment includes:
[0037] If the video memory capacity required for training a deep neural network is greater than the provided video memory capacity, then:
[0038] Perform the following operations before training:
[0039] A1. Set the storage state of each tensor of the deep neural network to the state of being written into the video memory, and denote the set composed of all tensors of the deep neural network as the tensor set;
[0040] A2. Based on the data dependency relationships among tensors in the tensor set, calculate the recomputation time for each tensor in the tensor set; obtain the tensor with the largest eviction time in the tensor set and the tensor with the largest ratio of tensor size to its recomputation time. If the generation cost of the former is less than that of the latter, set the storage state of the former to the memory swap state and remove it from the tensor set; otherwise, set the storage state of the latter to the recomputation state and remove it from the tensor set.
[0041] Since two factors are considered during sorting, one is the tensor size (the occupancy of video memory) and the other is the computation time (the overall training time), but the sorting methods for these two influencing factors are different for the two methods of recomputation and memory swap. Therefore, when making a selection in the present invention, a unified comparison standard is required, that is, to select according to the time to regenerate the tensor during backpropagation, so as to achieve the purpose of minimizing memory occupancy as much as possible under the premise of spending very little time regenerating the tensor. For example, if the computation time of a tensor is very short but its size is very small, the cost performance of memory swapping or recomputation for it is not high, because the main purpose of the present invention is to reduce memory occupancy, and secondly, to minimize the training time while reducing memory occupancy.
[0042] It should be noted that there can be various representation methods for the written-to-video-memory state, recomputation state, and memory swap state. In this embodiment, the written-to-video-memory state is represented by the storage flag bit 0; the recomputation state is represented by the storage flag bit 1; and the memory swap state is represented by the storage flag bit 2.
[0043] Preferably, in an alternative embodiment, as Figure 1 shown, the above step A2 includes:
[0044] Based on the data dependency relationships among tensors in the tensor set, calculate the recomputation time for each tensor in the tensor set; calculate the ratio of the size of each tensor in the tensor set to its recomputation time to obtain the memory released per unit time corresponding to each tensor.
[0045] Sort the tensors in the tensor set in descending order of their eviction times to obtain a memory swap priority sequence; sort the tensors in the tensor set in descending order of the memory released per unit time corresponding to them to obtain a recomputation priority sequence.
[0046] Compare the generation overheads of the first tensors in the memory swap priority sequence and the recomputation priority sequence during the backpropagation stage, select a storage scheme with a smaller overhead for the tensors, and modify the corresponding storage status; specifically, obtain the first tensors in the memory swap priority sequence and the recomputation priority sequence, and compare their generation overheads. If the former is less than the latter, set the storage status of the former to the memory swap status and remove it from the tensor set; otherwise, set the storage status of the latter to the recomputation status and remove it from the tensor set.
[0047] Among them, the generation overhead of the former is the ratio of the tensor size to the PCIe bandwidth; the generation overhead of the latter is the recomputation time of the tensor. Since the generation overhead of recomputation is related to the source tensor of the tensor, when its source tensor selects the memory swap storage method, the source tensor of the tensor will become the source tensor of the source tensor of the tensor, and it is necessary to reorder the overheads of the recomputation methods for this part of the tensors and re - make decisions for the tensors that have already been selected for recomputation.
[0048] It should be noted that since the subsequent sorting order of the memory swap priority sequence will not change after the first sorting is completed, only the first tensor in the memory swap priority sequence needs to be removed from the memory swap priority sequence. Therefore, the above - mentioned sorting method is the most efficient when obtaining the tensor with the largest eviction time and the tensor with the largest ratio of its size to its recomputation time, and can be used as a preferred implementation.
[0049] A3. Update the data dependency relationships of the tensors in the tensor set, and repeat steps A2 - A3 for iteration until the required video memory capacity for training the deep neural network is less than or equal to the provided video memory capacity; specifically, each time a memory swap method or a recomputation method is selected, since the tensors will be evicted during training and the data dependency relationships have changed, it is necessary to reorder the overheads of the recomputation methods. When the size of the tensor removed from the tensor set is greater than the missing memory size of the training model, stop further selection of evicted tensors and prepare to start training;
[0050] Furthermore, when a certain tensor A is removed from the tensor set, the tensors that depend on this tensor A directly depend on the tensors that tensor A depends on; when reflected in the tensor computation graph, it means that the child nodes of tensor A are directly pointed to the parent nodes of tensor A to update the data dependency relationships of the tensors in the tensor set.
[0051] During the training process: for tensors with the storage state of being written to the video memory, store them in the video memory during the forward propagation process; for tensors with the storage state of memory swapping, transfer them to the main memory during the forward propagation process and prefetch them from the main memory to the video memory before accessing the tensors during the backward propagation; for tensors with the storage state of recomputation, do not store the tensors during the forward propagation process, but recompute the tensors in the video memory according to the data dependency relationship before accessing the tensors during the backward propagation;
[0052] Among them, as Figure 2 shown, the eviction time of a tensor is the time interval between the forward propagation access and the backward propagation access of the tensor minus the transmission time of the tensor when training a deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network; the transmission time of a tensor is the sum of the time to unload the tensor to the main memory and the time to prefetch the tensor from the main memory when training a deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network; the recomputation time of a tensor is obtained by adding up the computation times of the tensors it depends on based on the data dependency relationship in the tensor set; the generation overhead of a tensor is the smaller value of the memory swapping overhead and the recomputation overhead of the tensor; the memory swapping overhead of a tensor is the ratio of the tensor size to the PCIe bandwidth; the recomputation overhead of a tensor is the recomputation time of the tensor.
[0053] Preferably, in an alternative implementation, the above memory management method further includes: step A0 executed before step A1, specifically including: training a deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network, and recording the access time (including the time accessed during the forward propagation of the tensor and the time accessed during the backward propagation), computation time, and size of each tensor. It should be noted that the training set used to train the deep neural network at this time is the same as the training set used in the subsequent training process.
[0054] Through the above method, the technical problem that the coarse-grained storage decision scheme based on the layer type of recomputation and memory swapping in the existing work limits the optimization space of memory usage can be effectively solved; specifically, before training, estimate the eviction time of the tensor and the memory released per unit time according to the size of the tensor and the parameters of the hardware, obtain the performance priority sequence of the recomputation and memory swapping methods of the tensor, and select the scheme with the smallest time overhead for the tensor based on the tensor granularity, while also improving the memory utilization rate.
[0055] During the training process, since memory swapping will cause memory fragmentation problems, the blocks released by the evicted tensors may not be able to be allocated and used during the subsequent training process. To further solve the technical problem of memory fragmentation caused by memory swapping, as Figure 3As shown, in an alternative embodiment, different placement strategies are implemented for tensors with a storage state of memory swap state (storage flag bit is 2) and tensors with a storage state of writing to video memory state (storage flag bit is 0), and they are placed in different storage blocks. By using different placement strategies for different types of tensors, the memory fragmentation problem caused by memory swapping is avoided; specifically, two storage blocks are allocated in the video memory, denoted as the default storage block and the swap storage block respectively, and their sizes can be the same or different. In this embodiment, the sizes of the default storage block and the swap storage block are the same;
[0056] During the forward propagation process, if the storage state of a tensor is the state of writing to video memory, the tensor is placed in the default storage block; if the storage state of the tensor is the memory swap state, the tensor is written into the swap storage block and transferred from the swap storage block to the main memory when the PCIe resource is idle; if the storage state of the tensor is the recomputation state, the tensor is not stored;
[0057] During the backward propagation process, if the storage state of a tensor is the memory swap state, the tensor is transferred to the swap storage block when the PCIe resource is idle; if the storage state of the tensor is the recomputation state, the tensor with the recomputation state after recomputation is placed in the default storage block.
[0058] Since the tensors with the storage state of memory swap state will soon be unloaded to the main memory, and the tensors prefetched to the video memory will also be released soon after being accessed, there are rarely situations of insufficient memory in the swap storage block. If such a situation occurs, only the size of the initially allocated swap storage block needs to be increased.
[0059] Generally speaking, in this embodiment, the eviction time of the tensor and the memory released per unit time are estimated according to the size of the tensor and the parameters of the hardware, and the scheme with the minimum generated overhead is selected for the tensor by means of overhead sorting. Memory management is carried out based on the tensor granularity, which improves the memory utilization rate. At the same time, the present invention further solves the memory fragmentation problem caused by the memory swapping method by using different placement strategies for different types of tensors, and further improves the memory utilization rate.
[0060] It should be noted that the above deep neural network can be a deep neural network used in scenarios such as image processing (such as image classification, object detection, etc.), natural speech processing, etc. When used for image processing, the above tensor is an image feature; when used for natural speech processing, the above tensor is a text feature.
[0061] Embodiment 2
[0062] A memory management system for a deep neural network, comprising: a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it executes the memory management method provided in Embodiment 1 of the present invention.
[0063] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.
[0064] Embodiment 3
[0065] A computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the memory management method provided in Embodiment 1 of the present invention.
[0066] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.
[0067] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A memory management method for deep neural networks, characterized in that Including: If the video memory capacity required for training a deep neural network is greater than the provided video memory capacity, then: Before training, perform the following operations: A1. Set the storage state of each tensor of the deep neural network to the state of being written to the video memory, and denote the set composed of all tensors of the deep neural network as the tensor set; A2. Based on the data dependence relationships among the tensors in the tensor set, calculate the recomputation time of each tensor in the tensor set; obtain the tensor with the largest eviction time in the tensor set and the tensor with the largest ratio of its size to its recomputation time. If the generation cost of the former is less than that of the latter, set the storage state of the former to the memory swap state and remove it from the tensor set; Otherwise, set the storage state of the latter to the recomputation state and remove it from the tensor set; A3. Update the data dependence relationships among the tensors in the tensor set, and repeat steps A2 - A3 for iteration until the video memory capacity required for training the deep neural network is less than or equal to the provided video memory capacity; During training: For a tensor with the storage state of being written to the video memory, store it in the video memory during the forward propagation process; For a tensor with the storage state of memory swap, transfer it to the main memory during the forward propagation process, and prefetch it from the main memory to the video memory before it is accessed during the backward propagation; For a tensor with the storage state of recomputation, do not store the tensor during the forward propagation process, but recompute the tensor in the video memory according to the data dependence relationships before it is accessed during the backward propagation; Wherein, the eviction time of a tensor is the time interval from when the tensor is accessed during the forward propagation to when it is accessed during the backward propagation minus the transmission time of the tensor when training the deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network; the generation cost of the former is the ratio of the tensor size to the PCIe bandwidth; the generation cost of the latter is the recomputation time of the tensor.
2. The memory management method according to claim 1, wherein The transmission time of a tensor is the sum of the time to unload the tensor to the main memory and the time to prefetch the tensor from the main memory when training the deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network.
3. The memory management method according to claim 1, wherein The said step A2 includes: Based on the data dependence relationships among the tensors in the tensor set, calculate the recomputation time of each tensor in the tensor set; calculate the ratio of the size of each tensor in the tensor set to its recomputation time to obtain the memory released per unit time corresponding to each tensor; Sort the tensors in the tensor set in descending order of their eviction times to obtain the memory swap priority sequence; sort the tensors in the tensor set in descending order of the memory released per unit time corresponding to them to obtain the recomputation priority sequence; Obtain the first tensors in the said memory swap priority sequence and the said recomputation priority sequence, and compare their generation costs. If the former is less than the latter, set the storage state of the former to the memory swap state and remove it from the tensor set; otherwise, set the storage state of the latter to the recomputation state and remove it from the tensor set.
4. The memory management method according to any one of claims 1-3, characterized in that The recomputation time of a tensor is obtained by summing up the computation times of the tensors it depends on, based on the data dependencies among tensors in the tensor set.
5. The memory management method according to any one of claims 1-3, characterized in that During the training process: Allocate two storage blocks in the video memory, denoted as the default storage block and the swap storage block respectively; For tensors with a storage state of being written to the video memory, store them in the default storage block during the forward propagation process; For tensors with a storage state of memory swapping, write them to the swap storage block during the forward propagation process, and transfer them from the swap storage block to the main memory when the PCIe resources are idle. Before accessing the tensors during the backward propagation, transfer them to the swap storage block when the PCIe resources are idle. For tensors with a storage state of recomputation, do not store the tensors during the forward propagation process, but recompute the tensors in the video memory according to the data dependencies before accessing the tensors during the backward propagation, and store them in the default storage block.
6. The memory management method according to any one of claims 1-3, characterized in that, It further includes: Step A0 executed before step A1, specifically including: training the deep neural network under the condition that the provided video memory capacity is greater than the video memory capacity required for training the deep neural network, and recording the access time, computation time, and size of each tensor.
7. The memory management method according to claim 1, characterized in that The deep neural network is a deep neural network for image processing or natural speech processing.
8. A memory management system for a deep neural network, characterized in that, It includes: A memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it executes the memory management method described in any one of claims 1 - 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the memory management method described in any one of claims 1 - 7.
Citation Information
Patent Citations
Pre-training model training processing method and device, electronic equipment and storage medium
CN114676761A
Method and device for training neural network
WO2020199914A1