Video memory management method, device, equipment and system

By dynamically adjusting the video memory resources according to task priority in the machine learning system, the problem of not being able to effectively guarantee the performance of high-priority tasks in the existing technology is solved, and higher resource utilization and task stability are achieved.

CN114443263BActive Publication Date: 2025-06-06ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011219652.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-03
Publication Date
2025-06-06
Estimated Expiration
2040-11-03

AI Technical Summary

Technical Problem

The prior art cannot effectively guarantee the performance of high-priority tasks when managing shared video memory resources of machine learning systems, and may lead to resource competition and task failure.

Method used

By determining the priority of multiple machine learning tasks running through the graphics processing unit, if the high-priority task requires video memory resources and the allocation resources are insufficient, the video memory resources occupied by the low-priority task are released to ensure that the high-priority task can be allocated to sufficient video memory space.

Benefits of technology

It realizes dynamic adjustment of video memory resources while ensuring the performance of high-priority tasks, improves the resource utilization rate of the overall cluster, and avoids task failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114443263B_ABST
    Figure CN114443263B_ABST
Patent Text Reader

Abstract

The present application discloses a video memory management method, device, system and equipment. The method includes: determining the priority of multiple machine learning tasks running through a graphics processing unit; if video memory resources are to be allocated to high-priority tasks, and the allocatable video memory resources are less than the video memory resource requirements of the high-priority tasks, then releasing at least a portion of the video memory resources occupied by low-priority tasks; allocating video memory resources to high-priority tasks to run high-priority tasks at least based on tensor data in the video memory space. By adopting this processing method, when the allocatable video memory resources are insufficient, the video memory resources occupied by low-priority tasks are allocated to high-priority tasks for use, thereby achieving dynamic scaling optimization of GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU, so that the resource utilization of the overall cluster can be improved while ensuring the performance of high-priority tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and specifically to a video memory management method, device and system, a machine learning system, and an electronic device. Background Art

[0002] With the continuous development of deep learning algorithms and the addition of graphics processing unit (GPU) computing power, deep learning has become a vital part of enterprise product data flow. In order to support large-scale deep learning applications, enterprises usually build shared GPU clusters to support the development of products across multiple fields, such as computer vision, natural language processing, speech recognition, recommendation and advertising services.

[0003] In order to improve the utilization of GPU resources and the throughput of the entire GPU cluster, the deep learning system allows multiple deep learning tasks to be run simultaneously on a GPU, so that more deep learning training tasks can be completed with the same amount of resources. At present, a typical way to reuse GPU memory resources is to manage the memory by a unified memory allocator within the deep learning framework. When the allocator receives a memory resource application from any task, as long as the GPU running the task has free memory resources, it will allocate the corresponding memory space to the task, regardless of the memory resource requirements of other tasks running on the GPU at the same time. This processing method can speed up the small batch training speed of tasks.

[0004] However, in the process of implementing the present invention, the inventors found that the above technical solutions all have at least the following problems: 1) The above resource reuse method does not provide any performance isolation guarantee, which will cause uncontrollable mutual influence between multiple tasks. Specifically, when the GPU is assigned to a "resource guarantee" task for separate use, the deep learning system can guarantee its task training performance. However, due to the lack of performance isolation mechanism on the GPU, if there are other tasks executed together on such a GPU, the potential competition for video memory resources may cause serious performance degradation of the "resource guarantee" task. 2) As the training progresses, the GPU video memory demand of the "resource guarantee" task may suddenly increase, and if the GPU video memory is occupied by other tasks at this time, the "resource guarantee" task will fail, which is even more unacceptable. In summary, how to manage the shared video memory resources of the machine learning system to improve the utilization of the GPU cluster while ensuring the use of video memory resources for high-priority tasks has become a problem that those skilled in the art urgently need to solve. Summary of the invention

[0005] The present application provides a video memory management method to solve the problem that the prior art cannot guarantee the performance of high-priority tasks. The present application also provides a video memory management device and system, a machine learning system, and an electronic device.

[0006] The present application provides a video memory management method, comprising:

[0007] Prioritize multiple machine learning tasks running through graphics processing units;

[0008] If video memory resources are to be allocated to the high-priority task, and the allocable video memory resources are less than the video memory resource requirement of the high-priority task, at least a portion of the video memory resources occupied by the low-priority task is released;

[0009] Allocate video memory resources to the high-priority task, so as to run the high-priority task at least according to the tensor data in the video memory space.

[0010] Optionally, also include:

[0011] Release the idle video memory resources occupied by the multiple machine learning tasks.

[0012] Optionally, releasing the idle video memory resources occupied by the multiple machine learning tasks includes:

[0013] Determine video memory resource usage status information of the machine learning task;

[0014] If the information satisfies the video memory resource release condition, the idle video memory resource is released.

[0015] Optionally, the usage status information includes: an upper limit value of the video memory resources actually used by the task;

[0016] The release condition includes: the duration for which the amount of video memory resource allocation of the task is greater than the upper limit value reaches a duration threshold.

[0017] Optionally, also include:

[0018] Release idle video memory resources for other high-priority tasks.

[0019] Optionally, releasing at least a portion of the video memory resources occupied by the low-priority task includes:

[0020] If the idle video memory resources of the low-priority task are greater than or equal to the required video memory resources, the idle video memory resources occupied by the low-priority task are released.

[0021] Optionally, releasing at least a portion of the video memory resources occupied by the low-priority task includes:

[0022] If the free video memory resources of the low priority task are less than the video memory resource demand, memory resources are allocated to the low priority task, and at least a portion of the video memory resources used by the low priority task are released, so as to continue running the low priority task at least according to the tensor data in the memory space.

[0023] Optionally, also include:

[0024] If the allocatable video memory resources increase to the video memory resource requirement of the low priority task, video memory resources are allocated to the low priority task to continue running the low priority task according to the tensor data in the video memory space.

[0025] Optionally, the low priority tasks include: iterative learning tasks;

[0026] The releasing of the video memory resources used by the low priority task includes:

[0027] After the low priority task completes the current iterative learning, at least a portion of the used video memory resources is released.

[0028] Optionally, before releasing at least a portion of the video memory resources used by the low-priority task, the method further includes:

[0029] Memory resources are allocated to high-priority tasks to run the high-priority tasks according to tensor data in the memory space and tensor data in the video memory space.

[0030] Optionally, after allocating video memory resources to the high-priority task, the method further includes:

[0031] Release memory resources for high priority tasks.

[0032] Optionally, the machine learning task includes a distributed deep learning task.

[0033] The present application also provides a video memory management method, including:

[0034] Running machine learning tasks on graphics processing units;

[0035] Determine video memory resource usage status information of the machine learning task;

[0036] If the information satisfies the video memory resource release condition, the idle video memory resources occupied by the task are released so that the idle video memory resources can be allocated to other machine learning tasks running in parallel through the graphics processing unit.

[0037] The present application also provides a video memory management device, comprising:

[0038] a prioritization unit for prioritizing a plurality of machine learning tasks to be run through the graphics processing unit;

[0039] A video memory releasing unit, configured to release at least a portion of the video memory resources occupied by the low priority task if video memory resources are to be allocated to the high priority task and the allocable video memory resources are less than the video memory resource requirement of the high priority task;

[0040] The video memory allocation unit is used to allocate video memory resources to high-priority tasks, so as to run the high-priority tasks at least according to the tensor data in the video memory space.

[0041] The present application also provides an electronic device, including:

[0042] Processor and memory;

[0043] A memory is used to store a program for implementing a video memory management method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: determining the priorities of multiple machine learning tasks run through a graphics processing unit; if video memory space is to be allocated to a high-priority task and the allocable video memory space is less than the video memory space requirement of the high-priority task, releasing at least a portion of the video memory space of the low-priority task; allocating video memory space to the high-priority task based on the video memory space released by the low-priority task, so as to run the high-priority task based on at least the tensor data in the video memory.

[0044] The present application also provides a video memory management device, comprising:

[0045] Task runners are used to run machine learning tasks using graphics processing units.

[0046] An information determination unit is used to determine the video memory resource usage status information of the machine learning task.

[0047] A video memory releasing unit is used to release the idle video memory resources occupied by the task if the information meets the video memory resource release condition, so as to allocate the idle video memory resources to other machine learning tasks running in parallel through the graphics processing unit.

[0048] The present application also provides an electronic device, including:

[0049] Processor and memory;

[0050] A memory for storing a program for implementing a video memory management method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: running a machine learning task through a graphics processing unit; determining video memory resource usage status information of the machine learning task; if the information meets a video memory resource release condition, releasing the idle video memory resources occupied by the task so that the idle video memory resources can be allocated to other machine learning tasks running in parallel through the graphics processing unit.

[0051] The present application also provides a video memory management system, including:

[0052] A storage resource coordinator is used to determine the priorities of multiple machine learning tasks running through the graphics processing unit; if video memory resources are to be allocated to the high-priority task and the allocable video memory resources are less than the amount of video memory resources of the high-priority task, a video memory resource release instruction is sent to the first storage resource allocator of the low-priority task, and a video memory resource allocation instruction is sent to the second storage resource allocator of the high-priority task;

[0053] A first storage resource allocator, used to release at least a portion of the video memory resources occupied by the low priority task according to the release instruction;

[0054] The second storage resource allocator is used to allocate video memory resources to the high priority task according to the allocation instruction, so as to run the high priority task at least according to the tensor data in the video memory space.

[0055] Optionally, the allocator is further used to send the video memory resource usage status information of the task to the coordinator;

[0056] The coordinator is further configured to send a video memory resource release instruction to the allocator if the information satisfies a video memory resource release condition.

[0057] Optionally, the distributor is specifically configured to send the information to the coordinator according to a preset period.

[0058] The present application also provides a machine learning system, comprising:

[0059] The client is used to send the priority information of the machine learning task to the server;

[0060] A server is used to determine the priorities of multiple machine learning tasks running through a graphics processing unit; if video memory resources are to be allocated to high-priority tasks and the allocatable video memory resources are less than the video memory resource requirements of the high-priority tasks, at least a portion of the video memory resources occupied by low-priority tasks are released; and video memory resources are allocated to high-priority tasks to run the high-priority tasks at least based on tensor data in the video memory space.

[0061] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned various methods.

[0062] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the above-mentioned various methods.

[0063] Compared with the prior art, this application has the following advantages:

[0064] The video memory management method provided in the embodiment of the present application determines the priorities of multiple machine learning tasks running through a graphics processing unit; if video memory resources are to be allocated to high-priority tasks and the allocatable video memory resources are less than the video memory resource demand of the high-priority tasks, at least a portion of the video memory resources occupied by low-priority tasks are released; video memory resources are allocated to high-priority tasks to run the high-priority tasks at least based on the tensor data in the video memory space; this processing method enables the video memory resources occupied by low-priority tasks to be allocated to high-priority tasks for use when the allocatable video memory resources are insufficient, thereby achieving dynamic scaling optimization of the GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU, so that the GPU video memory resources can be allocated to other tasks for use while ensuring the performance of high-priority tasks; therefore, the resource utilization of the entire cluster can be effectively improved while ensuring the performance of high-priority tasks.

[0065] The video memory management method provided in the embodiment of the present application runs a machine learning task through a graphics processing unit; determines the video memory resource usage status information of the machine learning task; if the information meets the video memory resource release condition, releases the idle video memory resources occupied by the task, so that the idle video memory resources can be allocated to other machine learning tasks running in parallel through the graphics processing unit; this processing method can release the idle video memory resources occupied by the task in a timely manner; therefore, it can effectively improve the resource utilization of the entire cluster.

[0066] The video memory management system provided by the embodiment of the present application determines the priorities of multiple machine learning tasks running through the graphics processing unit through a storage resource coordinator; if video memory resources are to be allocated to a high-priority task and the allocatable video memory resources are less than the video memory resource demand of the high-priority task, a video memory resource release instruction is sent to the first storage resource allocator of the low-priority task, and a video memory resource allocation instruction is sent to the second storage resource allocator of the high-priority task; the first storage resource allocator releases at least a portion of the video memory resources occupied by the low-priority task according to the release instruction; the second storage resource allocator allocates video memory resources to the high-priority task according to the allocation instruction, so as to run the high-priority task at least according to the tensor data in the video memory space; this processing method allows the local storage resource coordinator of the GPU device to allocate the video memory resources occupied by the low-priority task to the high-priority task for use when the allocatable video memory resources are insufficient, thereby realizing dynamic scaling optimization of the GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU, so that the GPU video memory resources can be allocated to other tasks for use while ensuring the performance of the high-priority task; therefore, the resource utilization of the overall cluster can be effectively improved while ensuring the performance of the high-priority task.

[0067] The machine learning system provided in the embodiment of the present application sends priority information of machine learning tasks to the server through the client; the server determines the priorities of multiple machine learning tasks running through the graphics processing unit; if video memory resources are to be allocated to high-priority tasks and the allocatable video memory resources are less than the video memory resource demand of the high-priority tasks, at least a portion of the video memory resources occupied by low-priority tasks are released; video memory resources are allocated to high-priority tasks to run the high-priority tasks at least based on the tensor data in the video memory space; this processing method enables the video memory resources occupied by low-priority tasks to be allocated to high-priority tasks for use when the allocatable video memory resources are insufficient, thereby achieving dynamic scaling optimization of the GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU, so that the GPU video memory resources can be allocated to other tasks for use while ensuring the performance of high-priority tasks; therefore, the resource utilization of the entire cluster can be effectively improved while ensuring the performance of high-priority tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 A flowchart of an embodiment of a video memory management method provided by the present application;

[0069] Figure 2 A schematic diagram of an application scenario of an embodiment of a video memory management method provided by the present application;

[0070] Figure 3 A schematic diagram of dynamic scaling of video memory resources in an embodiment of a video memory management method provided by the present application;

[0071] Figure 4 A schematic diagram of video memory resource changes in an embodiment of a video memory management method provided by the present application;

[0072] Figure 5 A schematic diagram of the structure of an embodiment of a video memory management device provided by the present application;

[0073] Figure 6 A schematic structural diagram of an embodiment of a video memory management system provided by the present application. DETAILED DESCRIPTION

[0074] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0075] In the present application, a video memory management method, device and system, a machine learning system, and an electronic device are provided. In the following embodiments, various solutions are described in detail one by one.

[0076] First embodiment

[0077] Please refer to Figure 1 , which is a flow chart of an embodiment of the video memory management method of the present application. The method provided in this embodiment may include the following steps:

[0078] Step S101: Determine the priorities of multiple machine learning tasks running through a graphics processing unit.

[0079] The video memory management method provided in this application can be applied in a machine learning system to allocate and use GPU video memory resources when multiple machine learning tasks share GPU video memory resources. The machine learning system can be a deep learning system built on a deep learning computing framework, such as TensorFlow, PyTorch and other deep learning computing frameworks. The machine learning task, also known as a machine learning model training task, can learn a machine learning model from training data. The model can be a model based on a deep neural network, and accordingly, the machine learning task is a deep learning task. For example, the model is a named entity recognition model, a speech recognition model, a product recommendation model, etc. learned from the training data. The model can also be a non-neural network machine learning model such as a decision tree.

[0080] In order to more intuitively illustrate the video memory management method provided by this application, the following first describes the application scenario of the method. Figure 2 As shown, the machine learning system can be a distributed machine learning system, including one or more GPU-based computing nodes, also known as GPU devices. Each GPU device may include one or more GPUs. Multiple machine learning tasks can be run in parallel on a GPU, and these machine learning tasks share the GPU's video memory resources. In addition, the GPU device also includes a central processing unit (CPU) and memory, wherein the CPU can also be called the GPU host. Figure 2 In the example, node 1 includes GPU1 and GPU2. Task D and Task E are run on GPU1 at the same time, and Task A, Task B, and Task C are run on GPU2 at the same time.

[0081] The video memory management method provided in the present application can dynamically adjust the video memory resources occupied by each task according to the performance guarantee priority (referred to as priority) of multiple tasks running in parallel on a GPU. The higher the priority of the task, the more its performance needs to be guaranteed. The priority of the task can be determined according to application requirements, such as setting only two priorities: high priority and low priority. For example, learning task 1 of the named entity recognition model is a "performance guarantee task" with a service level guarantee, and learning task 2 of the speech recognition model is a "speculative execution task" without a service level guarantee. In this case, task 1 can be set to a high priority and task 2 can be set to a low priority.

[0082] In specific implementation, multiple priorities may also be set. Table 1 shows the priority setting information of a machine learning task in an example.

[0083]

[0084]

[0085] Table 1. Machine learning task list

[0086] As can be seen from Table 1, three task priorities are set in this embodiment, where level 1 is a high priority task, level 2 is a second highest priority task, and level 3 is the lowest priority task. In this case, if there is a level 1 task in the parallel task, the performance of the level 1 task must be guaranteed first; if the level 1 task is not included, the performance of the level 2 task must be guaranteed.

[0087] Step S103: if video memory resources are to be allocated to the high priority task and the allocatable video memory resources are less than the video memory resource requirement of the high priority task, at least a portion of the video memory space occupied by the low priority task is released.

[0088] The video memory management method provided in the present application, when allocating video memory resources to a high-priority task, if the video memory resources available on the GPU are less than the video memory resource requirements of the high-priority task, part or all of the video memory space occupied by the low-priority task shall be released, and the released video memory space shall be allocated to the high-priority task, so that the tensor data of the high-priority task can be stored in the video memory as much as possible, thereby ensuring the running performance of the high-priority task.

[0089] The video memory resource requirement may be the video memory resources to be added by the deep learning task being run, or the video memory resources initially required by the deep learning task to be run. The video memory resource requirement may be determined by the deep learning task. Since the method for determining the video memory resource requirement is a relatively mature prior art, it will not be described here.

[0090] In this embodiment, if a deep learning task in the machine learning task table of a GPU is to be run, the video memory resources required by the task are first determined. After determining the video memory resources required for a task, if the allocatable video memory resources of the GPU are less than the video memory resource requirements of the task, the priority of the task and the priorities of other tasks running in the GPU are determined. If the priority of the task to be run is higher than the priority of other tasks running, in order to ensure the performance of the high-priority task, the video memory resources occupied by the low-priority task need to be released.

[0091] In specific implementation, the amount of video memory resources to be released can be determined based on the video memory resource requirements of the task to be run. For example, if a high-priority task requires 1G video memory resources, 1.2G of video memory resources occupied by a low-priority task can be released.

[0092] In specific implementation, the video memory resources occupied by one or more low-priority tasks can be released. For example, if a high-priority task requires 5G video memory resources, but releasing the video memory resources occupied by one low-priority task is not enough, the video memory resources occupied by multiple low-priority tasks can be released to release more than 5G video memory resources.

[0093] In one example, the method may further include the following step: releasing the idle video memory resources occupied by the multiple machine learning tasks. In the process of implementing the present invention, the inventors found that most deep learning tasks cannot fully utilize all GPU video memory resources allocated to them at all times, and there are usually idle video memory resources. After research, the inventors found that the idle video memory resources of deep learning tasks may be caused by the following reasons.

[0094] 1) Product-oriented deep learning training tasks usually include many computing parts, some of which are not easy to parallelize and therefore cannot fully occupy GPU memory resources, such as graph sampling in neural networks, feature extraction in advertising models, and data enhancement in computer vision.

[0095] 2) Deep learning systems face massive amounts of training data. As the amount of data increases, a lot of time is spent on network data synchronization in large-scale distributed training. Figure 2 Task E runs on both node 1 and node n-1, and the corresponding GPU memory resources are also idle during the data synchronization period.

[0096] 3) Distributed deep learning training usually uses the synchronous stochastic gradient descent (SGD) method, which requires that the resources required by the training task must be met at the same time, so that the training task can start. Therefore, from the perspective of the cluster scheduler, when resources are insufficient, the scheduler needs to reserve some available resources for distributed tasks until all required resources are met. This reservation process also causes the GPU memory resources to be in an idle waiting state.

[0097] 4) In practical applications, some tensor data are only used in specific deep learning training stages, such as data processing and model evaluation. After this stage, these tensor data will no longer be used and will be cleared from the video memory, thus generating idle video memory resources.

[0098] However, in the existing deep learning framework, the above-mentioned idle video memory resources will not be released. Instead, these video memory resources will always be reserved for the task. The overall video memory resources occupied by the task have not been reduced, but only a part of them is idle video memory. The task may use this part of available video memory in the subsequent running process, such as storing other tensor data, or it may not use this part of idle video memory resources. As a result, GPU resources are often still in a relatively low utilization state.

[0099] The method provided in this embodiment releases idle video memory resources that have not been fully utilized by tasks based on the above-mentioned video memory resource usage status, so that they can be promptly allocated to other tasks for use, thereby optimizing the processing of multiple tasks in a shared GPU resource scenario and avoiding queuing for other tasks, thereby improving GPU resource utilization and further improving the throughput of the shared GPU cluster.

[0100] In specific implementation, the release of idle video memory resources occupied by the multiple machine learning tasks may include the following sub-steps: 1) determining the video memory resource usage status information of the machine learning task; 2) if the information satisfies the video memory resource release condition, releasing the idle video memory resources. The usage status information includes: the upper limit of the video memory resources actually used by the task; the release condition includes: the duration of time when the video memory resource allocation amount of the task is greater than the upper limit and reaches a duration threshold. For example, during the running of the task, the upper limit (peak value) of the actual use of video memory resources is determined every 10 seconds. When the video memory resources occupied by the task are greater than the peak value actually required for 30 consecutive seconds (after determining the usage status information 3 times), the idle video memory resources will be released.

[0101] In one example, the releasing of at least a portion of the video memory space occupied by the low-priority task may include the following sub-steps: if the idle video memory resources of the low-priority task are greater than or equal to the video memory resource demand, then releasing the idle video memory resources occupied by the low-priority task. In this case, both types of tasks can be run based on the data stored in the video memory space, which not only ensures the running performance of the high-priority task, but also avoids affecting the performance of the low-priority task.

[0102] In another example, the releasing of at least a portion of the video memory space occupied by the low-priority task may also include the following sub-steps: if the free video memory resources of the low-priority task are less than the video memory resource demand, allocating memory resources to the low-priority task, and releasing at least a portion of the video memory resources used by the low-priority task, so as to continue running the low-priority task at least based on the tensor data in the memory space.

[0103] The video memory resources used by the low-priority task do not belong to the idle video memory resources, and the tensor data of the low-priority task is still stored in this part of the video memory resources. The low-priority task needs this resource to ensure the running performance. However, since part or all of the video memory resources used by the low-priority task are released, and part or all of the tensor data of the low-priority task is temporarily switched from the video memory space to the memory space for storage, the low-priority task will at least continue to run according to the tensor data in the memory space. If only a part of the video memory resources used by the low-priority task is released, the low-priority task can continue to run the low-priority task based on a part of the tensor data in the memory space and another part of the tensor data in the video memory space at the same time. If all the video memory resources used by the low-priority task are released, the low-priority task can continue to run based on all the tensor data in the memory space. It can be seen that no matter what the situation is, it will not cause the task to fail.

[0104] With this processing method, if the video memory resource demand of the high-priority task cannot be met after releasing the idle video memory resources of the low-priority task, it is necessary to further release the video memory resources occupied by the tensor data of the low-priority task, and transfer all or part of the tensor data of the low-priority task to the memory space of the GPU host machine for storage, at least run the low-priority task based on the tensor data in the memory space, so as to avoid the low-priority task and the high-priority task from competing for video memory resources, thereby ensuring the performance of the high-priority task, at the cost of sacrificing part of the performance of the low-priority task.

[0105] like Figure 3 As shown in the figure, sub-figure a shows the video memory resources occupied by a deep learning task in the initial stage, including the data of a tensor, where the height of the dotted line indicates the water level of the video memory resources occupied. Sub-figure b shows that as the tensor data required by the task increases accordingly, the upper limit of video memory usage also increases. These tensor data can be cached in the pool of the memory allocator (video memory resource allocator) of the task, that is, cached in the video memory space allocated for the task, so that the next small batch of tensor data of the task can continue to be reused. The increased video memory resources of the task can accommodate three tensor data. As the task continues to run, some tensor data are only used in a specific deep learning training stage, and after this stage, these tensor data will no longer be used, and these data will be cleared from the video memory, resulting in idle video memory resources. Sub-figure c shows that after clearing a tensor, idle video memory resources appear, and the video memory resource water level does not change.

[0106] The method provided in the embodiment of the present application dynamically adjusts the upper limit of video memory usage according to the video memory resources actually required by the task. In this embodiment, the video memory resources currently in use are actively detected, and the idle video memory resources are released, so as to adjust the upper limit of video memory usage to an appropriate value. Sub-figure d shows the video memory resource level after the above-mentioned available video memory (idle video memory resources) is recycled (released) after the available video memory has not been used for a period of time.

[0107] In this embodiment, the deep learning system exposes an interface to the outside world, allowing the upper limit of the GPU memory usage of the task to be increased or reduced while the task is running, and can even be reduced to a value lower than the actual demand of the task. Subgraph e can be a situation where when a high-priority task comes, the allocable video memory is insufficient, and the data of a low-priority tensor is transferred to the memory space of the host machine, so that the released video memory resources can be allocated to the high-priority task to ensure the performance of the high-priority task. With this processing method, even if there are not enough video memory resources, the low-priority task can continue to run and will not fail.

[0108] In one example, the method may further include the following steps: if the allocatable video memory resources grow to the video memory resource requirement of the low-priority task, then allocating video memory resources to the low-priority task to continue running the low-priority task based on the tensor data in the video memory space. With this processing method, when the GPU video memory resources are no longer in short supply, the video memory upper limit of the low-priority task can be increased, and the tensors of the low-priority task can also be reallocated on the GPU. Sub-figure f shows that when the GPU video memory resources are no longer in short supply, the tensor data of the low-priority task is re-stored from the memory space to the video memory space to maximize the performance of the low-priority task.

[0109] In one example, the low-priority task includes an iterative learning task; the releasing of the video memory resources used by the low-priority task can be implemented in the following manner: after the low-priority task completes the current iterative learning, the used video memory resources are released.

[0110] In deep learning scenarios, to perform the gradient descent algorithm on the training set, the entire data set is usually divided into several small training sets, and a small subset is trained each time. In this way, on the one hand, the huge amount of computation caused by the participation of the entire data set in training at one time can be avoided; on the other hand, the gradient direction of the subset will not be too different from that of the entire data set, which can ensure the correctness of the training. This training method is also called small batch training. This learning task is called an iterative learning task. Each time a small batch training is performed, various data involved in the training must be stored in the storage space. The container for storing these data is called a tensor. A tensor is a data unit for storing data in a deep learning framework. As a data container, it can store data during the training process. A task can include multiple tensors in one training. Tensor data can be an arbitrary number of multi-dimensional arrays composed of a set of primitive values, including training data for a small batch training, model parameters generated during the training process, and data of various intermediate nodes in the network.

[0111] In this embodiment, by moving tensor data between video memory and internal memory, the utilization rate of GPU video memory resources is improved while ensuring the running performance of high-priority tasks. The processing method of moving tensor data between video memory and internal memory allows internal memory resources to be used as video memory resources when video memory resources are insufficient, but this usually introduces huge data copy overhead.

[0112] The inventors of the present invention studied the unique characteristics of deep learning tasks, namely the characteristics of iterative learning, and found that in deep learning tasks, training is performed in small batches, tensor data will be created and destroyed within a small batch, and the same tensor will be repeatedly created between small batches. Therefore, it is chosen to dynamically adjust the upper limit of the video memory usage of the task on the boundary of the small batch. At this time, the tensor data has been released, so that explicit data copying between the video memory and the main memory can be avoided, thereby avoiding the huge copying overhead.

[0113] In a specific implementation, before releasing the video memory resources used by the low-priority task, the method may further include the following steps: allocating a portion of memory resources and a portion of video memory resources to the high-priority task, so as to run the high-priority task according to the tensor data in the memory space and the tensor data in the video memory space.

[0114] In a specific implementation, after allocating video memory resources for the high-priority task, the method may further include the following steps: releasing the memory resources of the high-priority task, so that the high-priority task can be run completely based on the tensor data in the video memory space.

[0115] like Figure 4As shown in a, the high-priority task suddenly increases its demand for video memory at time T0, and the video memory demand is a. However, the available video memory is not enough at this time, so it first occupies a in the memory space and temporarily runs the high-priority task based on the tensor data in the memory. At time T1, after a small batch training of the low-priority task is completed, the tensor data of the low-priority task is moved to the memory, and the video memory resource a released by the low-priority task is allocated to the high-priority task. At the same time, the memory resource a occupied by the high-priority task is released, so that the video memory resource of the high-priority task increases by a, and the high-priority task continues to be executed on the video memory. This processing method only reduces the performance of the high-priority task for a very short time, but does not cause failure, and can avoid huge copying overhead.

[0116] like Figure 4 As shown in b, during the time period from T0 to T1, the video memory resource b used by the low-priority task is released, and b is occupied in the memory space. The low-priority task runs based on the tensor data in the memory space, which loses some running performance. After that, after other tasks release the video memory resources, the video memory resources of the low-priority task at time T1 increase by b, and the low-priority task continues to run based on the tensor data in the video memory space.

[0117] In one example, not only the video memory space occupied by low-priority tasks can be released, but also the idle video memory space occupied by other high-priority tasks can be released. This processing method can release more video memory resources and effectively improve the performance of low-priority tasks.

[0118] It can be seen from the above embodiments that the video memory management method provided in the embodiments of the present application determines the priorities of multiple machine learning tasks running through a graphics processing unit; if video memory space is to be allocated to a high-priority task and the allocatable video memory space is less than the video memory space requirement of the high-priority task, at least a portion of the video memory space of the low-priority task is released; based on the video memory space released by the low-priority task, video memory space is allocated to the high-priority task to run the high-priority task at least based on the tensor data in the video memory; this processing method enables the video memory resources occupied by the low-priority task to be allocated to the high-priority task for use when the allocatable video memory resources are insufficient, thereby realizing dynamic scaling and optimization of the GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU, so that the GPU video memory resources can be allocated to other tasks while ensuring the performance of the high-priority task; therefore, the resource utilization of the overall cluster can be effectively improved while ensuring the performance of the high-priority task.

[0119] Second embodiment

[0120] In the above-mentioned embodiment, a video memory management method is provided. Correspondingly, the present application also provides a video memory management device. The device corresponds to the above-mentioned method embodiment. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0121] Please refer to Figure 5 , which is a schematic diagram of the structure of an embodiment of a video memory management device of the present application. The present application further provides a video memory management device, comprising:

[0122] a prioritization unit for prioritizing a plurality of machine learning tasks to be run through the graphics processing unit;

[0123] A video memory releasing unit, configured to release at least a portion of the video memory resources occupied by the low priority task if video memory resources are to be allocated to the high priority task and the allocable video memory resources are less than the video memory resource requirement of the high priority task;

[0124] The video memory allocation unit is used to allocate video memory resources to high-priority tasks, so as to run the high-priority tasks at least according to the tensor data in the video memory space.

[0125] Third embodiment

[0126] In the above-mentioned embodiment, a video memory management method is provided, and correspondingly, the present application also provides an electronic device. The device corresponds to the embodiment of the above-mentioned method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0127] An electronic device of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a video memory management method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining the priorities of multiple machine learning tasks run through a graphics processing unit; if video memory space is to be allocated to a high-priority task and the allocable video memory space is less than the video memory space requirement of the high-priority task, releasing at least a portion of the video memory space of the low-priority task; allocating video memory space to the high-priority task based on the video memory space released by the low-priority task, so as to run the high-priority task based on at least the tensor data in the video memory.

[0128] Fourth embodiment

[0129] Corresponding to the above-mentioned video memory management method, the present application also provides a video memory management method. The parts of this embodiment that are the same as those of the first embodiment will not be repeated here, please refer to the corresponding parts in the first embodiment.

[0130] In this embodiment, the method may include the following steps:

[0131] Step 1: Run the machine learning task via the graphics processing unit.

[0132] Step 2: Determine the video memory resource usage status information of the machine learning task.

[0133] Step 3: If the information satisfies the video memory resource release condition, the idle video memory resources occupied by the task are released so that the idle video memory resources can be allocated to other machine learning tasks running in parallel through the graphics processing unit.

[0134] The machine learning tasks include, but are not limited to, deep learning tasks. The idle video memory resources may include at least one of the following resources: idle video memory resources generated by modules that cannot be processed in parallel in deep learning tasks, and idle video memory resources generated when multiple resources required by deep learning tasks are being met.

[0135] The machine learning task may be a distributed deep learning task. The idle video memory resources may be idle video memory resources generated when multiple image processing units corresponding to the distributed deep learning task synchronize data.

[0136] Specifically, the idle video memory resources of deep learning tasks may be caused by the following reasons.

[0137] 1) Idle video memory resources generated by modules that cannot be processed in parallel in deep learning tasks. Product-oriented deep learning training tasks usually include many computing parts, some of which are not easy to parallelize and therefore difficult to fully occupy GPU video memory resources, such as graph sampling in neural networks, feature extraction in advertising models, and data enhancement in computer vision.

[0138] 2) Idle video memory resources generated when multiple image processing units corresponding to distributed deep learning tasks synchronize data. Deep learning systems face massive amounts of training data. As the amount of data increases, a lot of time is spent on the network data synchronization stage of the model in ultra-large-scale distributed training, such as Figure 2 Task E runs on both node 1 and node n-1, and the corresponding GPU memory resources are also idle during the data synchronization period.

[0139] 3) Idle video memory resources generated when waiting for multiple resources required by deep learning tasks to be met. Distributed deep learning training usually uses the synchronous stochastic gradient descent (SGD) method, which requires that the resources required by the training task must be met at the same time, so that the training task can start. Therefore, from the perspective of the cluster scheduler, when resources are insufficient, the scheduler needs to reserve some available resources for distributed tasks until all required resources are met. This reservation process also causes the GPU video memory resources to be in an idle waiting state.

[0140] 4) In practical applications, some tensor data are only used in specific deep learning training stages, such as data processing and model evaluation. After this stage, these tensor data will no longer be used and will be cleared from the video memory, thus generating idle video memory resources.

[0141] The usage status information includes, but is not limited to, the upper limit of the actual usage of the video memory resources of the task; the release condition includes, but is not limited to, the duration of the video memory resource allocation of the task greater than the upper limit reaches a duration threshold. The duration threshold can be determined according to application requirements, such as being set to 30 seconds.

[0142] For example, during the execution of a task, the upper limit value (peak value) of the actual usage of video memory resources is determined every 10 seconds. If the video memory resources occupied by the task are greater than the peak value actually required for 30 consecutive seconds (after determining the usage status information three times), the idle video memory resources will be released.

[0143] It can be seen from the above embodiments that the video memory management method provided in the embodiments of the present application runs a machine learning task through a graphics processing unit; determines the video memory resource usage status information of the machine learning task; if the information meets the video memory resource release conditions, releases the idle video memory resources occupied by the task, so that the idle video memory resources can be allocated to other machine learning tasks running in parallel through the graphics processing unit; this processing method enables the idle video memory resources occupied by the task to be released in a timely manner; therefore, the resource utilization of the overall cluster can be effectively improved.

[0144] Fifth embodiment

[0145] In the above-mentioned embodiment, a video memory management method is provided. Correspondingly, the present application also provides a video memory management device. The device corresponds to the above-mentioned method embodiment. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0146] In this embodiment, the device includes:

[0147] Task runners are used to run machine learning tasks using graphics processing units.

[0148] An information determination unit is used to determine the video memory resource usage status information of the machine learning task.

[0149] A video memory releasing unit is used to release the idle video memory resources occupied by the task if the information meets the video memory resource release condition, so as to allocate the idle video memory resources to other machine learning tasks running in parallel through the graphics processing unit.

[0150] Sixth embodiment

[0151] In the above-mentioned embodiment, a video memory management method is provided, and correspondingly, the present application also provides an electronic device. The device corresponds to the embodiment of the above-mentioned method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0152] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a video memory management method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: running a machine learning task through a graphics processing unit; determining video memory resource usage status information of the machine learning task; if the information meets a video memory resource release condition, releasing the idle video memory resources occupied by the task so that the idle video memory resources can be allocated to other machine learning tasks running in parallel through the graphics processing unit.

[0153] Seventh embodiment

[0154] Corresponding to the above-mentioned video memory management method, the present application also provides a video memory management system. The parts of this embodiment that are the same as those of the first embodiment will not be described in detail, please refer to the corresponding parts in the first embodiment.

[0155] Please refer to Figure 6 , which is a structural diagram of an embodiment of a video memory management device of the present application. A video memory management system provided by the present application includes: a storage resource coordinator, a storage resource allocator (first storage resource allocator) for low priority tasks (such as task A), and a storage resource allocator (second storage resource allocator) for high priority tasks (such as task B).

[0156] The system provided in this embodiment realizes adaptive dynamic scaling optimization of GPU memory resources for machine learning tasks through the collaborative design of the storage resource coordinator and the machine learning computing framework. Figure 6As shown, the storage resource coordinator supports adaptive GPU memory resource adjustment and can be deployed in a GPU computing node to schedule and manage the memory resources of one or more GPUs in the node, and dynamically adjust the usage of GPU memory resources for multiple tasks on a GPU. The machine learning computing framework corresponds to the machine learning task, and each task runs through the machine learning computing framework.

[0157] The machine learning computing framework can be an end-to-end machine learning platform. Various machine learning frameworks can have their own ecosystems, which include various tools, libraries and other resources, which can help developers easily build and deploy applications powered by machine learning. In this embodiment, the machine learning computing framework is a deep learning computing framework, including but not limited to TensorFlow, PyTorch, MXNet, Caffe, etc.

[0158] Among them, the storage resource coordinator is used to determine the priorities of multiple machine learning tasks running through the graphics processing unit; if video memory resources are to be allocated to high-priority tasks and the allocable video memory resources are less than the amount of video memory resources of the high-priority tasks, a video memory resource release instruction is sent to the storage resource allocator of the low-priority tasks, and a video memory resource allocation instruction is sent to the storage resource allocator of the high-priority tasks (the second storage resource allocator); the storage resource allocator of the low-priority tasks is used to release at least a portion of the video memory resources occupied by the low-priority tasks according to the release instruction; the storage resource allocator of the high-priority tasks is used to allocate video memory resources to the high-priority tasks according to the allocation instruction, so as to run the high-priority tasks at least based on the tensor data in the video memory space.

[0159] In one example, the allocator is further configured to send the video memory resource usage status information of the task to the coordinator. In a specific implementation, the allocator may send the information to the coordinator according to a preset period. Accordingly, the coordinator is further configured to send a video memory resource release instruction to the allocator if the information meets the video memory resource release condition.

[0160] It can be seen from the above embodiments that the video memory management system provided by the embodiments of the present application determines the priorities of multiple machine learning tasks running through the graphics processing unit through a storage resource coordinator; if video memory resources are to be allocated to high-priority tasks, and the allocable video memory resources are less than the video memory resource requirements of the high-priority tasks, a video memory resource release instruction is sent to the first storage resource allocator of the low-priority tasks, and a video memory resource allocation instruction is sent to the second storage resource allocator of the high-priority tasks; the first storage resource allocator releases at least a portion of the video memory resources occupied by the low-priority tasks according to the release instruction; the second storage resource allocator allocates at least a portion of the video memory resources occupied by the low-priority tasks according to the allocation instruction. High-priority tasks are allocated video memory resources to run high-priority tasks at least according to the tensor data in the video memory space. This processing method allows the local storage resource coordinator of the GPU device to allocate the video memory resources occupied by low-priority tasks to high-priority tasks when the allocable video memory resources are insufficient, thereby achieving dynamic scaling optimization of the GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU. In this way, GPU video memory resources can be allocated to other tasks while ensuring the performance of high-priority tasks. Therefore, the resource utilization of the entire cluster can be effectively improved while ensuring the performance of high-priority tasks.

[0161] Eighth embodiment

[0162] Corresponding to the above-mentioned video memory management method, the present application also provides a machine learning system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment. A machine learning system provided by the present application includes: a client and a server.

[0163] The client is used to send priority information of machine learning tasks to the server; the server is used to determine the priorities of multiple machine learning tasks run by the graphics processing unit; if video memory resources are to be allocated to high-priority tasks and the allocable video memory resources are less than the video memory resource requirements of the high-priority tasks, at least a portion of the video memory resources occupied by low-priority tasks are released; and video memory resources are allocated to high-priority tasks to run the high-priority tasks at least based on tensor data in the video memory space.

[0164] The client includes but is not limited to mobile communication devices, namely, commonly known as mobile phones or smart phones, and also includes terminal devices such as personal computers, PADs, iPads, etc. The server can run machine learning tasks on a GPU cluster.

[0165] In this embodiment, a task service device can be provided to the user through the client. The user can use the task service device of the client to determine the machine learning task to be run and set the task priority. For example, if the "secondary" priority is selected, different service fees may be required for different priorities. After determining the machine learning task to be run and setting the task priority, the task running request can be submitted to the server through the client. In response to the request, the server can store the priority information of the task and store the task in the task table.

[0166] In this embodiment, the server may include one or more GPU computing nodes, each of which may run a machine learning task through a machine learning framework, and a storage resource coordinator deployed in the GPU computing node may obtain priority information of multiple tasks to be run through the computing node from the server. Table 2 shows the machine learning task information in this embodiment.

[0167] Task ID User ID Task Priority Task 1 (Named Entity Recognition Model) User A Level 1 (Performance Assurance Mission) Task 2 (Named Entity Recognition Model) User B Level 2 (speculative execution tasks) Task 3 (Product Recommendation Model) User C Level 2 (speculative execution tasks) Task 4 (Language Model) User C Level 1 (Performance Assurance Mission) …

[0168] Table 2. Machine learning task list

[0169] As can be seen from Table 2, the learning task 1 of the named entity recognition model of user A is the “performance guarantee task”, and the learning task 2 of the named entity recognition model of user B is the “speculative execution task”. The priority of the “performance guarantee task” is higher than that of the “speculative execution task”.

[0170] After determining the priorities of multiple machine learning tasks running through the graphics processing unit, the following processing can be performed: if video memory resources are to be allocated to high-priority tasks and the allocatable video memory resources are less than the video memory resource requirements of the high-priority tasks, then at least a portion of the video memory resources occupied by the low-priority tasks are released; video memory resources are allocated to high-priority tasks to run the high-priority tasks at least based on the tensor data in the video memory space.

[0171] In one example, the server can also be used to determine performance information of a machine learning task; and adjust the priority of the task based on the performance information.

[0172] For example, the priority of task A was originally level 2, but the system did not enable the performance of the task to reach the "level 2" service level required by the user. In this case, the priority of the task can be adjusted to level 1 so that its actual performance can reach the "level 2" service level required by the user.

[0173] In specific implementation, the server can record the change information of video memory resources during the operation of machine learning tasks, and adjust the priority information of the task according to the change information. For example, if a high-priority task runs based on tensor data in memory space 30% of the time, and the performance information of the task does not meet the service level requirements, the priority of the task can be increased.

[0174] In another example, the server can also be used to determine performance information of a machine learning task; and determine serviceable machine learning tasks based on the performance information.

[0175] For example, the service level requirements of Task A and Task B can be met through the system, but the service level requirement of Task C cannot be met. In this case, video memory resource management services can be provided for Task A and Task B.

[0176] It can be seen from the above embodiments that the machine learning system provided by the embodiments of the present application sends priority information of machine learning tasks to the server through the client; the server determines the priorities of multiple machine learning tasks running through the graphics processing unit; if video memory resources are to be allocated to high-priority tasks and the allocatable video memory resources are less than the video memory resource demand of the high-priority tasks, at least a portion of the video memory resources occupied by the low-priority tasks are released; video memory resources are allocated to high-priority tasks to run the high-priority tasks at least based on the tensor data in the video memory space; this processing method enables the video memory resources occupied by low-priority tasks to be allocated to high-priority tasks for use when the allocatable video memory resources are insufficient, thereby achieving dynamic scaling optimization of the GPU video memory resources occupied by multiple machine learning tasks running in parallel on a GPU, so that the GPU video memory resources can be allocated to other tasks for use while ensuring the performance of high-priority tasks; therefore, the resource utilization of the overall cluster can be effectively improved while ensuring the performance of high-priority tasks.

[0177] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

[0178] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0179] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0180] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.

[0181] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A video memory management method, It is characterized in that include: Prioritize multiple machine learning tasks running through graphics processing units; If it is necessary to allocate video memory resources to a high-priority task with an increasing demand for GPU video memory, and the allocable video memory resources are less than the video memory resource demand of the high-priority task, and the idle video memory resources of the low-priority task are less than the video memory resource demand, and the low-priority task includes an iterative learning task, memory resources are allocated to the high-priority task to run the high-priority task according to the tensor data in the memory space and the tensor data in the video memory space; After the low-priority task completes the current iterative learning, memory resources are allocated to the low-priority task, and at least a portion of the video memory resources occupied by the low-priority task are released, so as to continue to run the low-priority task at least according to the tensor data in the memory space; Allocate video memory resources for high-priority tasks, so as to run the high-priority tasks at least according to the tensor data in the video memory space; If the allocatable video memory resources increase to the video memory resource requirement of the low priority task, video memory resources are allocated to the low priority task to continue running the low priority task according to the tensor data in the video memory space.

2. The method according to claim 1, It is characterized in that Also includes: Release the idle video memory resources occupied by the multiple machine learning tasks.

3. The method according to claim 2, It is characterized in that The releasing the idle video memory resources occupied by the multiple machine learning tasks includes: Determine video memory resource usage status information of the machine learning task; If the information satisfies the video memory resource release condition, the idle video memory resource is released.

4. The method according to claim 3, It is characterized in that The usage status information includes: the upper limit of the video memory resources actually used by the task; The release condition includes: the duration for which the amount of video memory resource allocation of the task is greater than the upper limit value reaches a duration threshold.

5. The method according to claim 1, It is characterized in that Also includes: Release idle video memory resources for other high-priority tasks.

6. The method according to claim 1, It is characterized in that The releasing at least a portion of the video memory resources occupied by the low-priority task comprises: If the idle video memory resources of the low-priority task are greater than or equal to the required video memory resources, the idle video memory resources occupied by the low-priority task are released.

7. The method according to claim 1, It is characterized in that After allocating video memory resources to the high priority task, the method further includes: Release memory resources for high priority tasks.

8. The method according to claim 1, It is characterized in that The machine learning task includes a distributed deep learning task.

9. A video memory management device, It is characterized in that include: a prioritization unit for prioritizing a plurality of machine learning tasks to be run through the graphics processing unit; A video memory release unit, configured to allocate memory resources to the high-priority task if video memory resources are to be allocated to the high-priority task with increasing demand for GPU video memory, and the allocatable video memory resources are less than the video memory resource demand of the high-priority task, and the idle video memory resources of the low-priority task are less than the video memory resource demand, and the low-priority task includes an iterative learning task, so as to run the high-priority task according to the tensor data in the memory space and the tensor data in the video memory space; After the low-priority task completes the current iterative learning, memory resources are allocated to the low-priority task, and at least a portion of the video memory resources occupied by the low-priority task are released, so as to continue to run the low-priority task at least according to the tensor data in the memory space; A video memory allocation unit is used to allocate video memory resources to high-priority tasks so as to run the high-priority tasks at least according to the tensor data in the video memory space; if the allocatable video memory resources grow to the video memory resource demand of the low-priority tasks, then allocate video memory resources to the low-priority tasks so as to continue running the low-priority tasks according to the tensor data in the video memory space.

10. An electronic device, It is characterized in that include: Processor and memory; A memory for storing a program for implementing the video memory management method according to any one of claims 1 to 8, wherein the device is powered on and runs the program of the method through the processor.

11. A video memory management system, It is characterized in that include: A storage resource coordinator to prioritize multiple machine learning tasks running through the graphics processing units; If video memory resources are to be allocated to a high-priority task with an increasing demand for GPU video memory, and the allocatable video memory resources are less than the amount of video memory resources of the high-priority task, and the idle video memory resources of the low-priority task are less than the amount of video memory resources required, and the low-priority task includes an iterative learning task, memory resources are allocated to the high-priority task to run the high-priority task according to the tensor data in the memory space and the tensor data in the video memory space; after the low-priority task completes the current iterative learning, memory resources are allocated to the low-priority task, and a video memory resource release instruction is sent to the first storage resource allocator of the low-priority task, and a video memory resource allocation instruction is sent to the second storage resource allocator of the high-priority task; if the allocatable video memory resources grow to the amount of video memory resources required by the low-priority task, a video memory resource allocation instruction is sent to the first storage resource allocator of the low-priority task; A first storage resource allocator, configured to release at least a portion of the video memory resources occupied by the low-priority task according to the release instruction, so as to continue to run the low-priority task at least according to the tensor data in the memory space; and to allocate video memory resources to the low-priority task according to the allocation instruction, so as to continue to run the low-priority task according to the tensor data in the video memory space; The second storage resource allocator is used to allocate video memory resources to the high priority task according to the allocation instruction, so as to run the high priority task at least according to the tensor data in the video memory space.

12. The system according to claim 11, It is characterized in that The distributor is further used to send the video memory resource usage status information of the task to the coordinator; The coordinator is further configured to send a video memory resource release instruction to the allocator if the information satisfies a video memory resource release condition.

13. The system according to claim 12, It is characterized in that The distributor is specifically configured to send the information to the coordinator according to a preset period.

14. A machine learning system, It is characterized in that include: The client is used to send the priority information of the machine learning task to the server; A server for determining the priorities of a plurality of machine learning tasks to be run through a graphics processing unit; if memory resources are to be allocated to a high-priority task with an increasing demand for GPU memory, and the allocatable memory resources are less than the memory resource demand of the high-priority task, and the idle memory resources of the low-priority task are less than the memory resource demand, and the low-priority task includes an iterative learning task, memory resources are allocated to the high-priority task to run the high-priority task according to the tensor data in the memory space and the tensor data in the memory space; After the low-priority task completes the current iterative learning, memory resources are allocated to the low-priority task, and at least a portion of the video memory resources occupied by the low-priority task are released, so as to continue running the low-priority task at least based on the tensor data in the memory space; video memory resources are allocated to the high-priority task, so as to continue running the high-priority task at least based on the tensor data in the video memory space; if the allocatable video memory resources grow to the video memory resource demand of the low-priority task, video memory resources are allocated to the low-priority task, so as to continue running the low-priority task based on the tensor data in the video memory space.

Citation Information

Patent Citations

  • GPU memory scheduler and GPU memory preemption method using the same

    KR102086757B1

  • Slab based memory management for machine learning training

    US20190286991A1