Method for memory management and system and method for machine learning
By using dependency structures and non-multiplex detectors in machine learning systems, identifying and relieving unnecessary data object storage, the problem of GPU memory performance bottleneck in machine learning training is solved, and training performance and efficiency are improved.
Patent Information
- Application Number
- CN201910040704.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-09
- Filing Date
- 2019-01-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-01-16
AI Technical Summary
Machine learning is susceptible to performance bottlenecks associated with GPU memory when executed on a graphics processor, especially when the memory is still unnecessarily allocated to cause performance degradation when data objects are no longer needed.
By generating a dependency structure, identifying the dependencies between the task and the data object, determining that unnecessary data object allocation is released after the task is completed, using a non-multiplex detector to mark and freeing the data object storage that is no longer needed.
Effectively manage memory resources, reduce unnecessary memory allocation and migration, and improve the performance and efficiency of machine learning training.
Smart Images

Figure CN110135588B_ABST
Abstract
Description
Technical Field
[0001] One or more aspects of embodiments according to the present invention relate to memory management, and more particularly, to a system and method for managing memory for machine learning. Background Art
[0002] Machine learning (ML) can suffer from performance bottlenecks related to GPU memory when executed on a graphics processing unit (GPU). Therefore, performance may suffer if memory is still unnecessarily allocated to certain data objects when those data objects are no longer needed for the computation to be performed.
[0003] Therefore, there is a need for an improved memory management system and method. Summary of the invention
[0004] According to an embodiment of the present invention, a method for memory management is provided, comprising: generating a dependency structure, the dependency structure comprising one or more task identifiers and one or more data object identifiers, the dependency structure comprising a list of one or more dependencies on a first data object identifier among the one or more data object identifiers, the first dependency of the list identifying a first task having a first data object identified by the first data object identifier as input; determining a count, the count being the number of dependencies on the first data object identifier; determining that the first task has completed execution; decrementing the count by 1 based at least in part on determining that the first task has completed execution; determining that the count is less than a first threshold; and deallocating the first data object based at least in part on determining that the count is less than the first threshold.
[0005] In one embodiment, the first threshold is 1.
[0006] In one embodiment, the method includes determining a number of dependencies associated with the first task.
[0007] In one embodiment, the first task is a computational operation in a first layer of a neural network.
[0008] In one embodiment, the first data object is an activation in the first layer.
[0009] In one embodiment, the first task includes, during a backward pass: computing gradients in the activations; and computing gradients in weights.
[0010] In one embodiment, the first data object is an input gradient in the first layer.
[0011] In one embodiment, the first task includes, during backward computation: computing gradients in activations; and computing gradients in weights.
[0012] In one embodiment, the first data object is a weight gradient in the first layer.
[0013] In one embodiment, the first task includes performing an in-place update on weights corresponding to the weight gradients.
[0014] In one embodiment, the method includes: generating a list of zero or more pass-persistent data object identifiers, a first pass-persistent data object identifier of the zero or more pass-persistent data object identifiers identifying a first data object in a neural network; determining that a pass-through has been completed; and deallocating the first data object based on determining that the pass-through has been completed.
[0015] In one embodiment, the first data object is the activations of a first layer of a neural network.
[0016] In one embodiment, the method includes generating a list of zero or more training-persistent data object identifiers, a first training-persistent data object identifier identifying a first data object in a neural network, determining that training of the neural network is complete, and deallocating the first data object based on determining that training of the neural network is complete.
[0017] In one embodiment, the first data objects are weights in a first layer of a neural network.
[0018] According to an embodiment of the present invention, a system for machine learning is provided, the system comprising: a graphics processor and a memory connected to the graphics processor, the graphics processor being configured to: call a noreuse detector; and after calling the noreuse detector, start a graphics processor kernel, the noreuse detector being configured to: identify a first data object, the first data object having a persistence defined at least by one or more tasks that use the data object as input; generate a dependency structure, the dependency structure comprising a first data object identifier that identifies the first data object, and a first task among the one or more tasks that uses the data object as input; determine a count, the count being the number of dependencies on the first data object identifier; determine that the first task has completed execution; decrement the count by 1 based at least in part on determining that the first task has completed execution; determine that the count is less than a first threshold; and deallocate the first data object based at least in part on determining that the count is less than the first threshold.
[0019] In one embodiment, the first threshold is 1.
[0020] In one embodiment, the first task is a computational operation in a first layer of a neural network.
[0021] In one embodiment, the first data object is an activation in the first layer.
[0022] In one embodiment, the first task includes, during backward computation: computing gradients in the activations; and computing gradients in weights.
[0023] According to an embodiment of the present invention, a method for machine learning is provided, the method comprising: allocating memory for a first data object in a neural network; determining that the first data object has a lifetime defined at least by one or more tasks that have the first data object as input; determining that a last one of the one or more tasks that have the first data object as input has completed execution; and deallocating the first data object based on determining that the last one of the one or more tasks that have the first data object as input has completed execution and based on determining that the first data object has a lifetime defined at least by the one or more tasks that have the first data object as input.
[0024] In one embodiment, the method includes: allocating memory for a second data object in the neural network; determining that the second data object has a lifetime defined by completion of a reverse calculation; and deallocating the second data object upon completion of the reverse calculation based on determining that the second data object has a lifetime defined by completion of the reverse calculation.
[0025] In one embodiment, the method includes: allocating memory for a third data object in the neural network; determining that the third data object has a lifetime defined by completion of training of the neural network; and deallocating the third data object upon completion of training of the neural network based on determining that the third data object has a lifetime defined by completion of training of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] These and other features and advantages of the present invention will be appreciated and understood with reference to the specification, claims and drawings, in which:
[0027] Figure 1 is a flow chart according to an embodiment of the present invention.
[0028] Figure 2 is a diagram of a directed acyclic graph according to an embodiment of the present invention.
[0029] Figure 3 is a data flow diagram of a forward pass according to an embodiment of the present invention.
[0030] Figure 4 4 is a data flow chart of reverse calculation according to an embodiment of the present invention.
[0031] Figure 5A is a vector and a graph according to an embodiment of the present invention.
[0032] Figure 5B is a data flow diagram according to an embodiment of the present invention.
[0033] [Explanation of Symbols]
[0034] 100: GPU memory;
[0035] 105: Slow system memory;
[0036] 110: Machine learning execution engine;
[0037] 115: Machine Learning Memory Manager;
[0038] 120: non-multiplexed detector;
[0039] A, B, C, D, E: data objects;
[0040] dE: error gradient;
[0041] dW x : Weight gradient / original gradient;
[0042] dW y : gradient / input / weight gradient;
[0043] dX, dY: gradient;
[0044] dZ: loss gradient / input / input gradient / gradient;
[0045] E: error;
[0046] mW x : sliding mean;
[0047] mW x *: updated sliding average;
[0048] W x , W x *、W y *: weight;
[0049] W y : weights / input;
[0050] X: data object / input;
[0051] Y: output / input / activation;
[0052] Z: output / predicted output / input / data object;
[0053] Z*: True output. DETAILED DESCRIPTION
[0054] The detailed description described below in conjunction with the accompanying drawings is intended as an illustration of an exemplary embodiment of a system and method for managing memory for machine learning provided according to the present invention, and is not intended to represent the only form in which the present invention can be constructed or utilized. The description sets forth the features of the present invention in conjunction with the illustrated embodiments. However, it should be understood that the same or equivalent functions and structures may be implemented by different embodiments that are also intended to be included within the scope of the present invention. As shown elsewhere herein, the same element numbers are intended to indicate the same elements or features.
[0055] In some related art systems, memory is allocated for data objects during machine learning training, and the data objects are retained until they are released, and some of the data objects are cached in GPU memory. Once the GPU memory reaches its maximum capacity, the system migrates the data objects allocated on the GPU to the system memory at the operating system (OS) level with page granularity. This approach can result in a loss of performance.
[0056] For machine learning (ML) operations on GPUs, limited GPU memory can cause performance bottlenecks. Therefore, some embodiments provide GPUs with large memory for efficient machine learning training by using slower but larger memory together with fast GPU memory.
[0057] Figure 1 The overall flow of some embodiments of using fast GPU memory 100 and slow system memory 105 on the host to provide large memory for the GPU is shown. Such embodiments may include a machine learning execution engine 110, a machine learning memory manager 115, and a non-reuse detector 120. The machine learning execution engine 110 executes GPU code and accesses data objects by calling the machine learning memory manager 115. The machine learning memory manager 115 may be a slab-based user-level memory management engine that manages data objects accessed by the GPU code executed by the machine learning execution engine 110. The non-reuse detector 120 distinguishes data objects based on data object type (or "class", as discussed in more detail below) and marks data objects that do not need to be retained by checking the directed acyclic graph (DAG) of the neural network so that the machine learning memory manager 115 can deallocate these data objects. The machine learning execution engine 110 performs a call to the non-reuse detector 120 after executing each GPU code.
[0058] In some embodiments, the above-mentioned non-reuse detector exploits the characteristics of machine learning training and alleviates performance-critical inefficiencies caused by unnecessarily requesting large dynamic random access memory (DRAM) sizes during machine learning training on GPUs. The non-reuse detector identifies and marks non-reused data objects so that the machine learning memory manager can deallocate these data objects to reduce data object migration overhead.
[0059] As in Figure 2In the example shown, the non-multiplexed detector examines a directed acyclic graph (DAG) of a neural network to identify non-multiplexed data objects. Figure 2 For example, it is shown that after computing "B" and "C", computations "D" and "E" are ready to be performed. In this example, the non-multiplexing detector recognizes that there are no other computations (or "tasks") that require "A" and marks "A" as "non-multiplexed" so that "A" can be deallocated by the machine learning memory manager.
[0060] Figure 3 and Figure 4 An example showing how a non-multiplexed detector can be used for machine learning training. The computations of each layer of the neural network are performed sequentially: from left to right for forward extrapolation, and from right to left for backward extrapolation. Figure 3 The first layer in the x After calculating the output "Y", the second layer takes the input "Y" and the weight "W y ” to calculate the output “Z”. After calculating the predicted output “Z”, the loss function compares “Z” with the true output “Z*” and calculates the error “E”.
[0061] Reference Figure 4 , the backward calculation starts by computing the loss gradient (or "input gradient") "dZ" with input "Z" and error gradient "dE". Then, the next layer takes input "dZ", "Y" and "W y ” to calculate the gradients “dY” and “dW y When inputting "dW y ” becomes available, the following calculation is performed: Execute the weight “W y " is updated to "W y *”. When the gradient “dY” becomes available, similarly Figure 4 Reverse calculation of the leftmost layer.
[0062] In some embodiments, the non-multiplexing detector generates a dependency structure that includes: one or more task identifiers, each of which identifies a corresponding task (e.g., a task that computes an output “Y” during forward pass, a task that computes an input gradient “dZ” during backward pass, and a task that computes a gradient “dY” and a weight “W” in an activation “Y” during backward pass). y The gradient "dW" in y ” task); and one or more data object identifiers (e.g., identifiers of data objects used as input to the task, such as “X” and “W” x” (the task used to compute the output “Y” during the forward pass), or “Y” and “dZ” (the task used to compute the gradient “dY” and gradient “dW” during the backward pass). y The dependency structure may link a data object identifier to a task that requires the data object identifier as input. For example, the dependency structure may include a list of one or more dependencies on a first data object identifier (e.g., an identifier of "X"), and a first dependency of the list may identify a first task (e.g., a task that computes output "Y" during forward calculation) that has as input the data object ("X") identified by the first data object identifier. The non-multiplexing detector may count, for example, the number of dependencies on the first data object identifier, reduce the count by 1 only when the first task completes execution, and deallocate the first data object when the count reaches zero. In this way, the lifetime of the first data object is defined by (or at least defined by or at least partially defined by) the one or more tasks that have the first data object as input.
[0063] In this example, the non-multiplexed detector in some embodiments converts the output of each layer during forward extrapolation ( Figure 3 and Figure 4 ) are marked as "non-reuse": when using these data objects to calculate the gradient ( Figure 4 After the gradient "dX", "dY" and "dZ" in the backward calculation are calculated, these data objects are no longer used. For example, after the gradient "dZ" is calculated, the data object "Z" is no longer referenced. Since the value of "Z" at the next iteration (machine learning training to process the next data item) does not depend on the value of "Z" at the previous iteration, there is no need to maintain the value of "Z" for the next iteration. In some embodiments, the non-reuse detector maintains a list of data objects in this category or "inference persistable" data objects (i.e., activated), and marks the data objects as "non-reuse" at the end of the backward calculation so that the data objects are deallocated when the backward calculation is completed. The non-reuse detector may further maintain a list of "training persistable" data object identifiers, each of which may identify a data object that is only deallocated when training is completed (e.g., a weight, such as "W x ”).
[0064] In some embodiments, the non-multiplexing detector also uses the dependency structure to use the weight gradients ( Figure 4 "dW x ”, “dW y”) to update the weights and then mark these data objects as “non-reuse”. Figure 5A and Figure 5B An example of weight update using momentum update is shown. Similar to various weight update methods, momentum update maintains the weight gradient ( Figure 5B "dW x ”) is a running average ( Figure 5A mW x ”), and apply a sliding average to update the weights (in Figure 5B From "W x " is updated to "W x *”) instead of applying the original gradient. After the weights and sliding averages are calculated and updated in-place, the original gradient ( Figure 5B "dW x ”). Deallocating the weight gradients computed during backward pass and the outputs of forward pass can reduce memory pressure on GPU memory during machine learning training on the GPU, thereby reducing unnecessary migrations between system memory and GPU memory on the GPU.
[0065] The term "processing circuit" is used herein to refer to any combination of hardware, firmware, and software used to process data or digital signals. The processing circuit hardware may include, for example, an application specific integrated circuit (ASIC), a general or dedicated central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), and a programmable logic device (such as a field programmable gate array (FPGA)). In the processing circuit used herein, each function is performed by hardware configured (i.e., hardwired) to perform the function, or by more general hardware (such as a CPU) configured to execute instructions stored in a non-temporary storage medium. The processing circuit may be fabricated on a single printed circuit board (PCB) or distributed on several interconnected PCBs. The processing circuit may include other processing circuits; for example, the processing circuit may include two processing circuits interconnected on a PCB, namely an FPGA and a CPU.
[0066] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Therefore, without departing from the spirit and scope of the present inventive concept, the first element, first component, first region, first layer or first section described herein may be referred to as the second element, second component, second region, second layer or second section.
[0067] The terms used herein are only used to illustrate specific embodiments and are not intended to limit the inventive concept. As used herein, the terms "substantially", "about" and similar terms are used as approximate terms rather than as terms of degree, and are intended to take into account the inherent deviations of the measured or calculated values that will be recognized by a person of ordinary skill in the art. As used herein, the term "major component" refers to a component present in the composition or product in an amount greater than the amount of any other single component in the composition, polymer or product. In contrast, the term "primary component" refers to a component that constitutes at least 50% or more by weight of a composition, polymer or product. As used herein, the term "major portion" means at least half of each item when applied to multiple items.
[0068] As used herein, the singular forms "a and an" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should be further understood that the terms "comprises and / or comprising" used in this specification specify the presence of stated features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. When preceding a series of elements, such as "at least one of..." Expressions such as "of" and "used" modify the entire series of elements and do not modify individual elements in the series. In addition, "may" used when describing an embodiment of the inventive concept refers to "one or more embodiments of the inventive concept." In addition, the term "exemplary" is intended to refer to an example or illustration. As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize, utilizing, and utilized," "utilizing," and "utilized," respectively.
[0069] It should be understood that when an element or layer is referred to as being "on," "connected to," "coupled to," or "adjacent to" another element or layer, the element or layer may be directly on, directly connected to, directly coupled to, or directly adjacent to the other element or layer, or one or more intervening elements or layers may be present. In contrast, when an element or layer is referred to as being "directly on," "directly connected to," "directly coupled to," or "immediately adjacent to" another element or layer, there are no intervening elements or layers.
[0070] Although exemplary embodiments of systems and methods for managing memory for machine learning have been specifically described and illustrated herein, many modifications and variations will be apparent to those skilled in the art. Therefore, it should be understood that the systems and methods for managing memory for machine learning constructed in accordance with the principles of the present invention may be implemented in ways other than those specifically described herein. The present invention is also defined in the following claims and their equivalents.
Claims
1. A method for memory management during machine learning training, comprising: Allocating fast graphics processor memory space and slow system memory space to the first data object; generating a dependency structure, the dependency structure comprising one or more task identifiers and one or more data object identifiers, the dependency structure comprising a list of one or more dependencies on the first data object identifier of the one or more data object identifiers, a first dependency of the list identifying a first task having as input a first data object identified by the first data object identifier; determining a count, the count being the number of dependencies on the first data object identifier; Determining that the first task has been completed; decrementing the count by one based at least in part on determining that the first task has completed execution; determining that the count is less than a first threshold; and The fast graphics processor memory space or the slow system memory space allocated to the first data object is deallocated based at least in part on determining that the count is less than the first threshold. The method according to claim 1 , wherein the first threshold is 1. . 3 . The method of claim 1 , further comprising determining a number of dependencies associated with the first task.
4. The method of claim 1, wherein the first task is a computational operation in a first layer of a neural network. The method of claim 4 , wherein the first data object is an activation in the first layer.
6. The method of claim 5, wherein the first task comprises during backcasting: computing gradients in the activations; and Compute the gradients in the weights.
7. The method of claim 4, wherein the first data object is an input gradient in the first layer.
8. The method of claim 7, wherein the first task comprises during backcasting: Compute gradients in activations; and Compute the gradients in the weights.
9. The method of claim 4, wherein the first data object is a weight gradient in the first layer.
10. The method of claim 9, wherein the first task comprises performing an in-situ update on weights corresponding to the weight gradients.
11. The method according to claim 1, further comprising: generating a list of zero or more inferred persistent data object identifiers, a first inferred persistent data object identifier identifying a first data object in the neural network; Make sure the backcasting is complete; as well as The first data object is deallocated based on determining that the backward calculation is complete.
12. The method of claim 11, wherein the first data object is an activation of a first layer of a neural network.
13. The method according to claim 1, further comprising: generating a list of zero or more training persistable data object identifiers, a first training persistable data object identifier identifying a first data object in the neural network, determining that training of the neural network is complete, and The first data object is deallocated based on determining that training of the neural network is complete.
14. The method of claim 13, wherein the first data object is a weight in a first layer of the neural network.
15. A system for machine learning, the system comprising: Graphics processor, and a fast graphics processor memory connected to the graphics processor, The graphics processor is configured to: Allocating the fast graphics processor memory space and the slow system memory space to the first data object; Invoking a non-multiplexed detector; and After calling the non-multiplexing detector, starting the graphics processor kernel, The non-multiplexed detector is configured to: identifying the first data object, the first data object having a lifetime defined by one or more tasks that have the first data object as input; Generate a dependency structure, the dependency structure comprising: a first data object identifier identifying said first data object, and a first task among the one or more tasks that takes the first data object as input; determining a count, the count being the number of dependencies on the first data object identifier; Determining that the first task has been completed; decrementing the count by one based at least in part on determining that the first task has completed execution; determining that the count is less than a first threshold; and The fast graphics processor memory space or the slow system memory space allocated to the first data object is deallocated based at least in part on determining that the count is less than the first threshold. The system of claim 15 , wherein the first threshold is 1. 18 .
17. The system of claim 15, wherein the first task is a computational operation in a first layer of a neural network.
18. The system of claim 17, wherein the first data object is an activation in the first layer.
19. The system of claim 18, wherein the first task comprises during backcasting: computing gradients in the activations; and Compute the gradients in the weights.
20. A method for machine learning, the method comprising: Allocating fast graphics processor memory space and slow system memory space for a first data object in the neural network; determining that the first data object has a lifetime defined at least by one or more tasks that have the first data object as input; determining that a last one of the one or more tasks having the first data object as input has completed execution; as well as The fast graphics processor memory space or the slow system memory space allocated to the first data object is deallocated based on determining that the last of the one or more tasks that have the first data object as input has completed execution and based on determining that the first data object has a lifetime defined by at least the one or more tasks that have the first data object as input.
21. The method of claim 20, further comprising allocating the fast graphics processor memory space and the slow system memory space for a second data object in the neural network; determining that the second data object has a lifetime defined by back-calculated completion; and The fast graphics processor memory space or the slow system memory space allocated to the second data object upon completion of the back extrapolation is deallocated based on determining that the second data object has the lifetime defined by the completion of the back extrapolation.
22. The method of claim 21, further comprising allocating the fast graphics processor memory space and the slow system memory space for a third data object in the neural network; determining that the third data object has a lifetime defined by completion of training of the neural network; and The fast graphics processor memory space or the slow system memory space allocated to the third data object upon the completion of the training of the neural network is de-allocated based on determining that the third data object has the lifetime defined by the completion of the training of the neural network.
Citation Information
Patent Citations
Information processing device and information processing method
WO2017073000A1