Tensor management method, electronic equipment, storage medium and program product

By generating tensor unloading rules and dynamically managing tensor migration, the problem of video memory overflow in deep neural network models is solved, training efficiency and stability are improved, and video memory utilization is optimized.

CN120295801AActive Publication Date: 2025-07-11XIAMEN UNIV

Patent Information

Application Number
CN202510785826.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

As the deep neural network model deepens and batches increase, tensors occupy a large amount of GPU memory (video memory), resulting in video memory overflow and affecting training efficiency.

Method used

By obtaining the tensor life cycle and training progress, generating unloading rules, prioritizing the unloading of high-yield and low-overhead tensors to the target storage device, dynamically manage the migration of tensors between the video memory and the target storage device, and optimizing video memory utilization.

Benefits of technology

Reduce memory usage, reduce overflow, improve model training efficiency and stability, optimize resource utilization, and avoid memory overflow and data transmission delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295801A_ABST
    Figure CN120295801A_ABST
Patent Text Reader

Abstract

The invention provides a tensor management method, electronic equipment, a storage medium and a program product. The tensor management method comprises the following steps: in a target model training process, obtaining a training progress and a tensor unloading rule of a target model; the tensor unloading rule is generated according to an unloading income-overhead ratio determined by a tensor life cycle of the target model; obtaining a stage unloading rule of the training progress from the tensor unloading rule; acquiring a resident tensor of the target model in the video memory; determining a tensor to be unloaded from the resident tensors according to a stage unloading rule; and unloading the tensor to be unloaded to the target storage device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and in particular, to a tensor management method, an electronic device, a storage medium, and a program product. Background Art

[0002] A deep neural network (DNN) model is a neural network model that includes an input layer, multiple hidden layers, and an output layer. The DNN model has the advantages of strong expressive ability (able to model complex non-linear systems), strong adaptability, strong scalability, and being good at processing large-scale data. The DNN model is the basis for many AI applications such as computer vision systems, speech recognition systems, natural language processing systems, recommendation systems, and medical diagnosis systems. With the wide application of the DNN model, the training calculation efficiency has become a key issue for improving the user experience.

[0003] In the DNN model, a tensor is the most basic data representation form, similar to a multi-dimensional array or matrix. Whether it is the input data of the DNN model, model parameters (such as weights, biases, etc.), or intermediate calculation results (such as activation values, gradients, etc.), they all exist in the form of tensors. During the training process of the DNN model, a large number of tensor creations, reads, modifications, and writes are involved in each forward propagation and backward propagation. These operations are frequent and the data volume is large. Especially in large model or large batch training, the number and scale of tensors will increase significantly. Since the Graphics Processing Unit (GPU) can parallelly process a large number of matrix and tensor operations, the training process of the DNN model is mostly executed on the GPU, thereby improving the training speed.

[0004] However, as the DNN model deepens and the batch size increases, tensors such as activation values, intermediate results, and gradients need to continuously occupy a large amount of GPU memory (i.e., video memory). When there are too many tensors or their sizes are too large, it is easy to cause video memory overflow, making the training process of the DNN model unable to run normally and affecting the training efficiency of the DNN model. Summary of the Invention

[0005] The present disclosure provides a tensor management method, an electronic device, a storage medium, and a program product.

[0006] According to one aspect of the present disclosure, a tensor management method is provided, including: during the training process of the target model, obtaining the training progress of the target model and the tensor offloading rule; the tensor offloading rule is generated based on the offloading benefit-cost ratio determined according to the tensor life cycle of the target model; obtaining the stage offloading rule of the training progress from the tensor offloading rule; obtaining the resident tensors of the target model in the video memory; determining the tensors to be offloaded from the resident tensors according to the stage offloading rule; and offloading the tensors to be offloaded to the target storage device.

[0007] According to the technical solution of one aspect, offloading the resident tensors in the video memory through the tensor offloading rule can reduce the occupancy of the video memory, thereby alleviating the pressure on the video memory, further reducing video memory overflow and improving the training efficiency of the model.

[0008] According to the tensor management method of at least one embodiment of the present disclosure, before training the target model, it further includes: obtaining the tensors of the target model and each stage of the training process of the target model; respectively obtaining the tensor life cycle of each tensor of the target model; respectively determining the offloading benefit-cost ratio of each tensor in each stage according to the tensor life cycle of each tensor; and generating a tensor offloading rule according to the offloading benefit-cost ratio of each tensor in each stage.

[0009] According to the technical solution of this embodiment, it is possible to improve the stability of the target model training and provide data support for the allocation of video memory resources.

[0010] According to the tensor management method of at least one embodiment of the present disclosure, for any tensor of the target model, the step of respectively determining the offloading benefit-cost ratio of each tensor in each stage according to the tensor life cycle of each tensor includes: obtaining the tensor size and tensor offloading cost of the tensor; obtaining the tensor idle time of the tensor in each stage according to the tensor life cycle of the tensor; determining the tensor offloading benefit of the tensor in each stage according to the product of the tensor size and the tensor idle time of the tensor in each stage; and determining the offloading benefit-cost ratio of the tensor in each stage according to the quotient of the tensor offloading benefit and the tensor offloading cost of the tensor in each stage; optionally, the step of obtaining the tensor size and tensor offloading cost of the tensor includes: obtaining the tensor size and network bandwidth of the tensor; and determining the tensor offloading cost according to the quotient of twice the tensor size of the tensor and the network bandwidth.

[0011] According to the technical solution of this embodiment, it is possible to preferentially select tensors with a larger offloading benefit-cost ratio for offloading when offloading tensors, thereby reducing the number of tensor offloading times, further reducing the offloading cost, and improving the tensor management efficiency.

[0012] According to the tensor management method of at least one embodiment of the present disclosure, after unloading the tensor to be unloaded to the target storage device, it further includes: obtaining the tensor life cycle and the current time of the tensor to be unloaded; determining the remaining time according to the tensor life cycle and the current time of the tensor to be unloaded; judging whether the remaining time is greater than the time threshold; in response to the remaining time of the target tensor in the tensor to be unloaded not being greater than the time threshold, migrating the target tensor to the video memory; optionally, migrating the target tensor to the video memory includes: adding the target tensor to a prefetch queue; asynchronously taking out the tensors in the prefetch queue and then adding them to the video memory.

[0013] According to the technical solution of this embodiment, it can ensure that the target tensor reaches the video memory in time while optimizing the utilization of the video memory, without causing GPU computing waiting. Furthermore, in the case of video memory priority, it maximizes the computing performance of the video memory and avoids video memory overflow or data transmission delay.

[0014] According to the tensor management method of at least one embodiment of the present disclosure, before training the target model, it further includes: in response to receiving a model training task, adding the model training task to a first queue; obtaining the real-time remaining amount of the video memory and the video memory required for each model training task in the first queue; determining the tasks to be trained and the delayed training tasks from the model training tasks in the first queue according to the real-time remaining amount and the video memory required for each model training task in the first queue; migrating the delayed training tasks from the first queue to a second queue, and migrating the tasks to be trained from the first queue to a third queue; the second queue is used to manage the delayed training tasks and migrate the delayed training tasks back to the first queue when a first recall instruction is triggered; controlling the execution of the tasks to be trained in the third queue, and the tasks to be trained in the third queue include the task of training the target model.

[0015] According to the technical solution of this embodiment, it can reduce model training task conflicts, optimize load distribution, avoid resource overload, and improve the scalability and reliability of the tensor management system.

[0016] According to the tensor management method of at least one embodiment of the present disclosure, in the process of controlling the execution of the tasks to be trained in the third queue, it further includes: monitoring whether there is a risk of super-resolution in the video memory; in response to the video memory having a risk of super-resolution, obtaining the task migration benefit ratio of each task to be trained in the third queue; determining the tasks to be migrated from the tasks to be trained in the third queue according to the task migration benefit ratio; migrating the tasks to be migrated from the third queue to a fourth queue; the fourth queue is used to manage the tasks to be migrated and migrate the tasks to be migrated back to the third queue when a second recall instruction is triggered.

[0017] According to the technical solution of this embodiment, it is possible to avoid video memory super-resolution at a relatively small cost, thereby improving scheduling flexibility, avoiding system crashes, and maximizing the overall benefit when video memory resources are tight.

[0018] According to the tensor management method of at least one embodiment of the present disclosure, obtaining the task eviction benefit ratio of each to-be-trained task in the third queue includes: obtaining the video memory demand growth amount and the historical eviction times of each to-be-trained task in the third queue; adding 1 to the historical eviction times of each to-be-trained task in the third queue to obtain the expected swap-out times of each to-be-trained task in the third queue; determining the task eviction benefit ratio of the to-be-trained task according to the quotient of the video memory demand growth amount and the expected swap-out times of each to-be-trained task in the third queue; optionally, monitoring whether there is a risk of video memory super-resolution includes: obtaining the video memory demand of the to-be-trained tasks in the third queue and the real-time remaining amount of the video memory; obtaining the size relationship between the video memory demand and the real-time remaining amount; determining whether there is a risk of video memory super-resolution according to the size relationship; optionally, obtaining the video memory demand of the to-be-trained tasks in the third queue and the real-time remaining amount of the video memory includes: obtaining the total number of tensor elements, the tensor element size, and the fixed video memory overhead of the to-be-trained task; determining the tensor video memory overhead according to the product of the total number of tensor elements and the tensor element size; determining the video memory demand of the to-be-trained task according to the sum of the tensor video memory overhead and the fixed video memory overhead; optionally, determining the to-be-evicted task from the to-be-trained tasks in the third queue according to the task eviction benefit ratio includes: sorting the to-be-trained tasks in the third queue according to the high and low of the task eviction benefit ratio of the to-be-trained tasks in the third queue to obtain a task sorting result; determining the number of evicted tasks of the to-be-trained tasks in the third queue according to the task eviction rule; determining the to-be-evicted tasks with the number of evicted tasks from the task sorting result.

[0019] According to the technical solution of this embodiment, it is possible to preferentially evict the to-be-trained tasks with "high pressure, few swaps, and high eviction benefits", thereby reducing the risk of video memory super-resolution and avoiding problems such as the to-be-trained task not being executed due to frequent eviction of a certain to-be-trained task, resulting in training failure.

[0020] According to another aspect of the present disclosure, there is provided an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the tensor management method of any one of the embodiments of the present disclosure.

[0021] According to yet another aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, and when the execution instructions are executed by a processor, they are used to implement the tensor management method of any one of the embodiments of the present disclosure.

[0022] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the tensor management method according to any one of the embodiments of the present disclosure.

[0023] The tensor management method provided by the present disclosure unloads the resident tensors in the video memory according to the tensor offloading rules, reduces the occupancy of the video memory, thereby alleviating the pressure on the video memory, further reducing video memory overflow and improving the training efficiency of the model. This tensor management method solves the problem in the prior art that as the DNN model deepens and the batch size increases, tensors such as activation values, intermediate results, and gradients need to continuously occupy a large amount of GPU memory (i.e., video memory), and when there are too many tensors or their sizes are too large, it is easy to cause video memory overflow, making the training process of the DNN model unable to run normally and affecting the training efficiency of the DNN model. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.

[0025] Figure 1 is the flowchart of the tensor management method according to an embodiment of the present disclosure Figure 1 .

[0026] Figure 2 is the flowchart of the tensor management method according to an embodiment of the present disclosure Figure 2 .

[0027] Figure 3 is Figure 2 the flowchart of the calculation method of the offloading benefit - overhead ratio in the tensor management method shown.

[0028] Figure 4 is Figure 3 the flowchart of the method for obtaining the offloading overhead in the calculation method of the offloading benefit - overhead ratio shown.

[0029] Figure 5 is the flowchart of the tensor management method according to an embodiment of the present disclosure Figure 3 .

[0030] Figure 6 is Figure 5 the flowchart of the tensor detour method in the tensor management method shown.

[0031] Figure 7 is the flowchart of the tensor management method according to an embodiment of the present disclosure Figure 4 .

[0032] Figure 8Flowchart of the tensor management method according to an embodiment of the present disclosure Figure 5 。

[0033] Figure 9 is Figure 8 Flowchart of the calculation method of the task migration benefit ratio in the tensor management method shown in the figure.

[0034] Figure 10 is Figure 8 Flowchart of the video memory monitoring method in the tensor management method shown in the figure.

[0035] Figure 11 is Figure 9 Flowchart of the video memory requirement acquisition method in the video memory monitoring method shown in the figure.

[0036] Figure 12 is Figure 8 Flowchart of the task determination method in the tensor management method shown in the figure.

[0037] Figure 13 Schematic flowchart of the tensor management method according to an embodiment of the present disclosure.

[0038] Figure 14 is Figure 13 Schematic diagram of task scheduling in the tensor management method shown in the figure.

[0039] Figure 15 Schematic block diagram of the structure of the tensor management device according to an embodiment of the present disclosure.

[0040] Figure 16 Schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0041] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content, rather than limiting the present disclosure. Additionally, it should be noted that, for the sake of convenience of description, only parts related to the present disclosure are shown in the drawings.

[0042] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.

[0043] Suppose an image classification model based on ResNet-152 is being trained for an autonomous vehicle to sense and recognize traffic scenes, such as identifying road signs, traffic lights, pedestrians, and obstacles. This image classification model has approximately 60.2 million parameters, and the training data is high-definition RGB images with a resolution of 1024×1024. During training, taking Batch Size (batch size) as 64 as an example, only the input data occupies approximately 768MB of video memory, the parameters and gradients of the image classification model together occupy approximately 480MB of video memory, the intermediate activation values and optimizer states together occupy more than 2GB of video memory, and adding system overheads such as caches and temporary tensors, the overall video memory occupancy can reach 7 to 8GB. If one tries to increase the Batch Size to 128 to improve the training speed, the total video memory occupancy will exceed 12GB. For a GPU with a maximum available video memory capacity of 12GB or less, it is very easy to trigger a video memory overflow error, resulting in the interruption of the training process, and thus affecting the training efficiency of the image classification model.

[0044] To this end, the present disclosure proposes a tensor management method, an electronic device, a storage medium, and a program product, which can be implemented by tensor management software installed on an electronic device such as a server or a cloud computing platform with powerful parallel processing capabilities and video memory resources and capable of efficiently processing large-scale tensor computing tasks.

[0045] For the convenience of description and to make the technical solutions of the specific embodiments of the present disclosure easier to understand, the technical terms involved in the present disclosure are explained as follows: A tensor is a multi-dimensional array or matrix used to store and transfer data.

[0046] The tensor life cycle refers to the entire process from the creation to the destruction of a tensor, usually including an active period of the tensor during which the tensor is frequently used, updated, and computed, and an inactive period of the tensor during which the tensor is no longer frequently involved in computations.

[0047] Figure 1 The overall flowchart of the tensor management method M100 according to an embodiment of the present disclosure is shown. As Figure 1 The shown tensor management method includes steps S110 to S150.

[0048] Specifically, Figure 1 The shown tensor management method includes the following.

[0049] Step S110, during the training process of the target model, obtain the training progress of the target model and the tensor offloading rule.

[0050] In some embodiments of the present disclosure, the target model in step S110 is generally a model in the fields of deep learning and machine learning, especially a Deep Neural Networks (DNN) model (such as AlexNet, VGGNet, ResNet, etc.) and its variants. The training progress of the target model can be pre-divided into multiple stages based on different tasks and optimization objectives, such as the initialization stage, the training stage, the convergence stage, the overfitting stage, etc. The training progress obtained through step S110 can be one of the above pre-divided stages.

[0051] The tensor offloading rule in step S110 is generated based on the offloading benefit-cost ratio determined by the tensor life cycle of the target model. The offloading benefit-cost ratio is the ratio of the benefit to the cost of tensor offloading determined by the tensor life cycle. Among them, the benefit is the video memory saved after offloading the tensor. The larger the tensor and the longer the idle time, the greater the benefit; the cost is the latency of tensor offloading and prefetching.

[0052] Step S110 can obtain the training progress based on information such as a loss function or an evaluation metric. Among them, when obtaining the training progress based on the loss function, it can be obtained based on the changes in the training loss and the validation loss in the loss function; for example, when both the training loss and the validation loss are high and there is no obvious downward trend, the obtained training progress is the initialization stage; when both the training loss and the validation loss decrease, the obtained training progress is the training stage; when both the training loss and the validation loss are low and tend to be stable, the obtained training progress is the convergence stage; when the training loss continues to decrease while the validation loss begins to rise or stagnate, the obtained training progress is the overfitting stage.

[0053] When obtaining the training progress based on an evaluation metric, the evaluation metric can be accuracy, recall rate, etc. Taking accuracy as an example, when the accuracy is low, the obtained training progress is the initialization stage; when the accuracy gradually rises, the obtained training progress is the training stage; when the accuracy tends to be stable and no longer changes significantly, the obtained training progress is the convergence stage.

[0054] Step S120, obtain the stage offloading rule of the training progress from the tensor offloading rule.

[0055] In some embodiments of the present disclosure, step S120 can obtain the stage offloading rule from the tensor offloading rule using the training progress obtained through step S110 as a retrieval condition.

[0056] Step S130, obtain the resident tensors of the target model in the video memory.

[0057] In some embodiments of the present disclosure, the resident tensor obtained through step S130 is the tensor of the target model that is currently retained in the video memory and has not been released. Step S130 can obtain the resident tensor of the target model in the video memory by means of video memory usage analysis, tensor lifecycle analysis, etc.

[0058] Step S140, determine the tensors to be unloaded from the resident tensors according to the stage unloading rules.

[0059] In some embodiments of the present disclosure, the process of determining the tensors to be unloaded through step S140 may include: first, determine the unloading priority of each tensor in the resident tensors according to the stage unloading rules; then, determine a preset number or a preset proportion of the tensors to be unloaded from the resident tensors in descending order of the unloading priority.

[0060] Step S150, unload the tensors to be unloaded to the target storage device.

[0061] In some embodiments of the present disclosure, step S150 can unload the tensors to be unloaded by inserting an unloading instruction during the process of compiling the program for training the target model. The target storage device in step S150 can be the system memory; in particular, the target storage device can also be a solid state drive (SSD), etc. Among them, the system memory is suitable for tasks with frequent tensor access or computationally intensive tasks, and the SSD is suitable for tasks that require a large amount of storage space.

[0062] The tensor management method provided by the present disclosure unloads the resident tensors in the video memory according to the tensor unloading rules, reduces the occupancy of the video memory, thereby alleviating the pressure on the video memory, and further reducing video memory overflow and improving the training efficiency of the model. This tensor management method solves the problem in the prior art that as the DNN model deepens and the batch size increases, tensors such as activation values, intermediate results, and gradients need to continuously occupy a large amount of GPU memory (i.e., video memory), and when there are too many tensors or the size is too large, it is easy to cause video memory overflow, making the training process of the DNN model unable to run normally and affecting the training efficiency of the DNN model.

[0063] Furthermore, the tensor management method provided by the present disclosure, before training the target model (i.e., before step S110), may further include steps S160 to S190 as Figure 2 shown.

[0064] Step S160, obtain the tensors of the target model and each stage of the training process of the target model.

[0065] In some embodiments of the present disclosure, all tensors related to the target model can be obtained through step S160, such as input tensors for representing training data or validation data inputs, weight tensors for weighting and transforming input data, bias tensors for adjusting the output of each neuron, gradient tensors for updating weights and biases, etc.

[0066] Each stage of the training process of the target model obtained through step S160 includes the stage corresponding to the training progress in step S110, and each stage of the training process of the target model can be pre-divided according to the training tasks and optimization objectives of the target model.

[0067] Step S170: Obtain the tensor lifecycle of each tensor of the target model respectively.

[0068] In some embodiments of the present disclosure, step S170 can estimate the tensor lifecycle based on the computational graph of the target model and the execution time of the Graphics Processing Unit (GPU) kernel program. Specifically, the process of obtaining the tensor lifecycle through step S170 can include: obtaining the computational graph of the target model and the execution time of the GPU kernel program; determining the execution order of key operations related to the tensor lifecycle according to the computational graph; obtaining the time information of the key operations of each tensor of the target model according to the execution time of the GPU kernel program; constructing the timeline of the key operations of each tensor according to the execution order and time information; and determining the tensor lifecycle according to the timeline.

[0069] Step S180: Determine the offloading benefit-cost ratio of each tensor in each stage respectively according to the tensor lifecycle of each tensor.

[0070] In some embodiments of the present disclosure, the offloading benefit-cost ratio in step S180 is an indicator of the potential benefit of offloading a tensor from video memory to the target storage device, and can be used to evaluate whether a tensor is worth offloading.

[0071] Step S190: Generate a tensor offloading rule according to the offloading benefit-cost ratio of each tensor in each stage.

[0072] In some embodiments of the present disclosure, the tensor offloading rule in step S190 specifies which tensor to offload in which stage during the training process of the target model. For any stage, the way of generating the tensor offloading rule in step S190 can be to sort the offloading benefit-cost ratios of each tensor in this stage from high to low to generate the tensor offloading rule, and the tensor offloading rule can include offloading targets, offloading quantities (or offloading ratios), etc.

[0073] In particular, to improve the efficiency of tensor management, steps S160 to S190 can be executed offline during the compilation phase of the target model training program, so that no unloading decision is required during the target model training process, and only the tensor unloading rules need to be followed, thereby improving the stability of the target model training and providing data support for video memory resource allocation.

[0074] Steps S160 to S190 generate tensor unloading rules based on the unloading benefit-cost ratio, so that tensors with a relatively large unloading benefit-cost ratio are preferentially unloaded during tensor unloading, thereby enabling the video memory to effectively support the target model training, reducing video memory overflow, and improving resource utilization and training efficiency.

[0075] For any tensor of the target model, regarding step S180, in some embodiments of the present disclosure, it may include steps S181 to S184 as Figure 3 shown.

[0076] Step S181, obtain the tensor size and tensor unloading cost of the tensor.

[0077] In some embodiments of the present disclosure, the tensor size obtained through step S181 is the space size occupied by the tensor in the video memory, usually in bytes (B), KB, MB, GB; the tensor size is usually the product of the number of elements included in the tensor and the number of bytes occupied by a single element. The tensor unloading cost obtained through step S181 is the cost required for the tensor to be unloaded from the video memory to the target storage device, usually including the migration cost and the prefetch cost.

[0078] Step S182, obtain the tensor idle time of the tensor in each stage according to the tensor life cycle of the tensor.

[0079] In some embodiments of the present disclosure, the tensor idle time in step S182 is the time interval between the current time and the next time the tensor is used. The tensor is usually idle at the current time; if the tensor is occupied at the current time, the tensor idle time is specifically the time interval between the current use of the tensor and the next use.

[0080] Step S183, determine the tensor unloading benefit of the tensor in each stage according to the product of the tensor size and the tensor idle time of the tensor in each stage.

[0081] Step S184, determine the unloading benefit-cost ratio of the tensor in each stage according to the quotient of the tensor unloading benefit of the tensor in each stage and the tensor unloading cost.

[0082] Obtaining the tensor offloading benefit through steps S181 to S184 enables the prioritized selection of tensors with relatively large offloading benefit costs for offloading during tensor offloading, thereby reducing the number of tensor offloading times, further reducing the offloading cost, and improving the tensor management efficiency.

[0083] Regarding step S181, in some embodiments of the present disclosure, it may include steps S1811 to S1812 as Figure 4 shown.

[0084] Step S1811, obtain the tensor size and network bandwidth of the tensor.

[0085] In some embodiments of the present disclosure, the network bandwidth obtained through step S1811 is the upper limit of the data volume that the network can transmit per unit time. This network bandwidth can specifically be the theoretical network bandwidth or the actual network bandwidth.

[0086] Step S1812, determine the tensor offloading cost of the tensor according to the quotient of twice the tensor size of the tensor and the network bandwidth.

[0087] Steps S1811 to S1812 determine the tensor offloading cost based on the tensor size and network bandwidth, which can assist in the offloading decision.

[0088] Furthermore, for the tensor management method provided by the present disclosure, after step S150, it may further include steps S200 to S230 as Figure 5 shown.

[0089] Step S200, obtain the tensor life cycle and the current time of the tensor to be offloaded.

[0090] Step S210, determine the remaining time according to the tensor life cycle and the current time of the tensor to be offloaded.

[0091] In some embodiments of the present disclosure, the remaining time in step S210 is used to determine how long the tensor to be offloaded will be used. The process of determining the remaining time through step S210 may include: determining the next usage time of the tensor according to the tensor life cycle; determining the remaining time according to the difference between the next usage time of the tensor and the current time.

[0092] Step S220, determine whether the remaining time is greater than the time threshold.

[0093] In some embodiments of the present disclosure, the time threshold in step S220 may be preset, and this time threshold is usually related to the transmission bandwidth between the video memory and the target storage device.

[0094] When it is determined through step S220 that the remaining time of the target tensor in the tensor to be unloaded is not greater than the time threshold, step S230 is executed; otherwise, steps S200 to S220 are re-executed until the remaining time is not greater than the time threshold or the current training process of the target model ends.

[0095] Step S230: Migrate the target tensor to the video memory.

[0096] This tensor management method can dynamically migrate tensors between the video memory and the target storage device through steps S110 to S150 and steps S200 to S230. It can insert corresponding tensor unloading instructions and prefetch instructions at the compiler intermediate expression level generated during the compilation of the program for training the target model, realizing the automatic management of tensors during the training process of the target model. For example: insert a prefetch instruction before the instruction to access the tensor by load or store, and insert an evict instruction after the end of the tensor life cycle.

[0097] By controlling the time when the target tensor in the tensor to be unloaded migrates back to the video memory through steps S200 to S230, it can ensure that the target tensor arrives at the video memory in time while optimizing the utilization of the video memory, without causing GPU computing waiting. Furthermore, in the case of giving priority to the video memory, it can maximize the computing performance of the video memory and avoid video memory overflow or data transmission delay.

[0098] Regarding step S230, in some embodiments of the present disclosure, it may include steps S231 to S232 as Figure 6 shown.

[0099] Step S231: Add the target tensor to the prefetch queue.

[0100] In some embodiments of the present disclosure, the prefetch queue in step S231 is a structure for asynchronous tensor migration.

[0101] Step S232: Asynchronously take out the tensor from the prefetch queue and then add it to the video memory.

[0102] In some embodiments of the present disclosure, after migrating the tensor in the prefetch queue back to the video memory through step S232, the tensor position and the video memory mapping table of the tensor migrated back to the video memory can be synchronously updated for subsequent use.

[0103] Asynchronously migrating the tensor back to the video memory through steps S231 to S232 can make the tensor ready in the video memory when the tensor is needed, thus avoiding the delay of computing waiting for data loading. It can also reduce the I / O bottleneck, balance the I / O load, and thus improve the throughput.

[0104] Furthermore, the tensor management method provided by the present disclosure, before training the target model, may further include asFigure 7 Steps S240 to S280 shown above.

[0105] Step S240, in response to receiving a model training task, add the model training task to the first queue.

[0106] In some embodiments of the present disclosure, the first queue in step S240 is used to manage all received model training tasks, and the first queue can specifically be a waiting queue.

[0107] Step S250, obtain the real-time remaining amount of video memory and the video memory required for each model training task in the first queue.

[0108] In some embodiments of the present disclosure, the real-time remaining amount obtained through step S250 is the amount of memory in the video memory that has not been allocated or used; the video memory required for each model training task obtained through step S250 is the total amount of video memory required during the execution of the model training task.

[0109] Step S260, determine the tasks to be trained and the tasks with delayed training from the model training tasks in the first queue according to the real-time remaining amount and the video memory required for each model training task in the first queue.

[0110] In some embodiments of the present disclosure, step S260 can sort the video memory required for each model training task in the first queue from small to large, and determine whether the minimum required video memory is greater than the real-time remaining amount; if it is greater, all model training tasks in the current first queue are regarded as tasks with delayed training; if it is not greater, the model training task is regarded as a task to be trained, and the latest real-time remaining amount is updated (the latest real-time remaining amount = the current real-time remaining amount - the minimum required video memory), and the judgment process is executed again based on the latest real-time remaining amount until all model training tasks in the first queue are determined.

[0111] Step S270, migrate the tasks with delayed training from the first queue to the second queue, and migrate the tasks to be trained from the first queue to the third queue.

[0112] In some embodiments of the present disclosure, the tasks with delayed training indicate that the current video memory is not sufficient to accommodate all the computing requirements of the task, and executing it may cause video memory overflow. Therefore, it is necessary to migrate the tasks with delayed training from the first queue to the second queue through step S270.

[0113] The second queue in step S270 is used to manage the tasks with delayed training, and migrates the tasks with delayed training back to the first queue when the first recall instruction is triggered. The second queue can specifically be a delay queue. The second queue can adopt a first-in, first-out strategy to ensure that any model training task will not be delayed indefinitely.

[0114] After migrating the delayed training task from the first queue to the second queue through step S270, the model training task in the second queue can be migrated back to the first queue according to a preset first recovery rule. The first recovery rule can be based on one or more settings such as waiting time, video memory resources, retry window, etc.

[0115] In some embodiments of the present disclosure, the training task to be processed indicates that the current video memory is sufficient to accommodate the computing requirements of the task, and the training task to be processed can be executed smoothly. Therefore, the training task to be processed can be migrated from the first queue to the third queue through step S270.

[0116] In step S270, the third queue is used to control the execution of the training task to be processed, and the third queue can specifically be an execution queue.

[0117] Step S280, control the execution of the training task to be processed in the third queue. The training task to be processed in the third queue includes the task of training the target model.

[0118] Steps S240 to S280 manage the execution of the model training task through the first queue, the second queue, and the third queue, which can reduce model training task conflicts, optimize load distribution, avoid resource overload, and improve the scalability and reliability of the tensor management system.

[0119] Furthermore, in the process of controlling the execution of the training task to be processed in the third queue through step S280 in the tensor management method provided by the present disclosure, steps S290 to S320 as shown in Figure 8 can also be included.

[0120] Step S290, monitor whether there is a risk of over-scoring in the video memory.

[0121] In some embodiments of the present disclosure, the over-scoring risk in step S290 is the risk that the total amount of memory requested by the application program or system from the video memory exceeds the actual available physical capacity of the video memory. Step S290 can determine whether there is an over-scoring risk in the video memory based on continuously detecting the video memory occupancy rate; for example, a over-scoring threshold is preset. When the video memory occupancy rate is not greater than the over-scoring threshold, there is no over-scoring risk in the video memory, and when the video memory occupancy rate is greater than the over-scoring threshold, there is an over-scoring risk in the video memory.

[0122] Specifically, a time-related video memory demand change function can be pre-constructed, and it is determined whether there is an over-scoring risk in the video memory according to the video memory demand change function and the real-time remaining amount of the video memory. The video memory demand change function is: total video memory demand = total number of elements of the tensors in the video memory of the model to be trained * average data type size of the tensors in the video memory of the model to be trained + fixed additional overhead for training the model to be trained (such as framework overhead, etc.).

[0123] When it is monitored through step S290 that there is no risk of out-of-memory in the video memory, continue to execute the monitoring process through step S290 until the model training task in the video memory is completed; when it is monitored through step S290 that there is a risk of out-of-memory in the video memory, execute step S300.

[0124] Step S300, obtain the task relocation benefit ratio of each to-be-trained task in the third queue.

[0125] In some embodiments of the present disclosure, the task relocation benefit ratio in step S300 is a decision-making index for measuring whether it is worth relocating tasks.

[0126] Step S310, determine the to-be-relocated tasks from the to-be-trained tasks in the third queue according to the task relocation benefit ratio.

[0127] Step S320, relocate the to-be-relocated tasks from the third queue to the fourth queue.

[0128] In some embodiments of the present disclosure, the fourth queue in step S320 is used to manage the to-be-relocated tasks and relocate the to-be-relocated tasks back to the third queue when a second recall instruction is triggered. The fourth queue can be a pause queue.

[0129] It is possible to continuously select the to-be-relocated tasks with a relatively large task relocation benefit ratio from the second queue through steps S290 to S320 and relocate them to the fourth queue until it is predicted that the video memory usage will not exceed its total amount in the next period of time. After relocating the to-be-relocated tasks to the fourth queue, the memory occupied by the to-be-relocated tasks can be released, thereby avoiding possible out-of-memory situations. When there are tasks in the fourth queue, it is possible to preferentially schedule the tasks in the fourth queue rather than the newly migrated tasks in the third queue.

[0130] After relocating the to-be-relocated tasks from the third queue to the fourth queue through step S320, the to-be-relocated tasks in the fourth queue can be recalled back to the third queue according to a preset second recovery rule. The second recovery rule can be set according to one or more of the video memory resource situation, the priority of the to-be-relocated tasks, the time when the model training tasks are added to the fourth queue, the task rotation strategy, etc.

[0131] When there is a risk of out-of-memory in the video memory through steps S290 to S320, relocating the to-be-relocated tasks in the to-be-trained tasks in the third queue to the fourth queue according to the task relocation benefit ratio can avoid out-of-memory at a relatively small cost, thereby improving the scheduling flexibility, avoiding system crashes, and maximizing the overall benefit when the video memory resources are tight.

[0132] Steps S240 to S320 utilize the variation in the use of video memory resources during the life cycle of the model training task to mine potential video memory utilization and computing parallelism. By actively selecting model training tasks and moving them backward, resource allocation conflicts can be avoided, and the utilization rate of video memory and training efficiency can be improved.

[0133] Regarding step S300, in some embodiments of the present disclosure, it may include steps S301 to S303 as Figure 9 shown.

[0134] Step S301, obtain the growth amount of video memory demand and the historical migration times of each training task to be trained in the third queue.

[0135] In some embodiments of the present disclosure, the growth amount of video memory demand in step S301 is the amount of video memory that needs to be occupied for the execution of a certain training task to be trained; the historical migration times in step S301 is the number of times a certain training task to be trained has been migrated to the fourth queue before the current time.

[0136] Step S302, add 1 to the historical migration times of each training task to be trained in the third queue to obtain the expected swap-out times of each training task to be trained in the third queue.

[0137] Step S303, determine the task migration benefit ratio of each training task to be trained in the third queue according to the quotient of the growth amount of video memory demand and the expected swap-out times of each training task to be trained in the third queue.

[0138] Determining the task migration benefit ratio based on the growth amount of video memory demand and the expected swap-out times through steps S301 to S303 helps to preferentially migrate training tasks to be trained that are "under high pressure, have few swap-outs, and high migration benefits", thereby reducing the risk of video memory over-scoring and avoiding problems such as the training task not being executed due to frequent migration of a certain training task, resulting in training failure.

[0139] Regarding step S290, in some embodiments of the present disclosure, it may include steps S291 to S293 as Figure 10 shown.

[0140] Step S291, obtain the video memory demand of the training task to be trained in the third queue and the real-time remaining amount of the video memory.

[0141] Step S292, obtain the size relationship between the video memory demand and the real-time remaining amount.

[0142] Step S293, determine whether there is a risk of video memory over-scoring according to the size relationship.

[0143] In some embodiments of the present disclosure, when the size relationship obtained through step S292 is that the video memory requirement is greater than the real-time remaining amount, it can be determined through step S293 that there is a risk of super-resolution in the video memory; when the size relationship obtained through step S292 is that the video memory requirement is not greater than the real-time remaining amount, it can be determined through step S293 that there is no risk of super-resolution in the video memory.

[0144] Determining whether there is a risk of super-resolution in the video memory based on the video memory requirement and the real-time remaining amount through steps S291 to S293 is of great significance for ensuring the stability of DNN training and optimizing resource utilization.

[0145] Regarding step S291, in some embodiments of the present disclosure, it may include steps S2911 to S2913 as Figure 11 shown.

[0146] Step S2911: Obtain the total number of tensor elements, the size of tensor elements, and the fixed video memory overhead of the task to be trained.

[0147] In some embodiments of the present disclosure, the total number of tensor elements of the task to be trained in step S2911 is the total amount of all elements in the tensor of the task to be trained. The size of tensor elements in step S2911 is the number of bytes occupied by a single tensor element of the task to be trained in memory; specifically, the size of tensor elements can be the average value of the tensor element sizes of all tensors of the task to be trained. The fixed video memory overhead in step S2911 is the additional video memory required to support the training of the task to be trained except for the tensor video memory overhead; the fixed video memory overhead may include stable video memory consumption such as framework overhead, memory allocator overhead, library loading overhead, and driver overhead.

[0148] Step S2912: Determine the tensor video memory overhead according to the product of the total number of tensor elements and the size of tensor elements.

[0149] Step S2913: Determine the video memory requirement of the task to be trained according to the sum of the tensor video memory overhead and the fixed video memory overhead.

[0150] Based on the total number of tensor elements, the size of tensor elements, and the fixed video memory overhead, steps S2911 to S2913 can accurately estimate the video memory requirement, thereby preventing video memory overflow in advance and improving training efficiency.

[0151] Regarding step S310, in some embodiments of the present disclosure, it may include steps S311 to S313 as Figure 12 shown.

[0152] Step S311: Sort the tasks to be trained in the third queue according to the task migration benefit ratio of the tasks to be trained in the third queue to obtain a task sorting result.

[0153] Step S312: Determine the number of training tasks to be migrated out of the third queue according to the task migration rule.

[0154] In some embodiments of the present disclosure, the task migration rule in step S312 may be a predefined task migration ratio or task migration quantity. When the task migration rule is the task migration ratio, step S312 may specifically include: obtaining the total number of training tasks to be trained in the third queue; determining the number of migrated tasks based on the product of the total number and the task migration ratio (when the product of the total number and the task migration ratio is not an integer, the value obtained by rounding down or up the product of the total number and the task migration ratio may be used as the number of migrated tasks). When the task migration rule is the task migration quantity, step S312 may directly use the predefined task migration quantity as the number of training tasks to be migrated out of the third queue.

[0155] Step S313: Determine the tasks to be migrated out with the number of migrated tasks from the task sorting result.

[0156] In some embodiments of the present disclosure, step S313 may specifically be to determine the tasks to be migrated out with the number of migrated tasks in descending order of the task migration benefit ratio according to the task sorting result.

[0157] By steps S311 to S313, tasks with a high task migration benefit ratio can be preferentially migrated out, which can significantly relieve the video memory pressure on the premise of minimizing performance loss.

[0158] For the tensor management method provided by the present disclosure, since the access pattern of the tensors of the target model has regularity and predictability, the tensors in the video memory can be managed based on the tensor unloading rule determined by the tensor life cycle. During the training process of the target model, only a very small part of the tensors need to be retained in the video memory to complete the current calculation, and most of the remaining tensors can be unloaded to the target storage device and prefetched to the video memory in advance when needed, avoiding affecting the execution of the model training task, which can greatly reduce the video memory occupancy, improve the video memory utilization rate, and reduce the risk of video memory overflow. In addition, by optimizing the resource utilization rate of the video memory and the target storage device, it is also possible to reduce the dependence on high-end graphics card hardware and lower the system deployment cost.

[0159] Specifically, for most DNN models, since both the forward propagation and backward propagation of the DNN model are carried out layer by layer, and the current calculation of each layer of the network only involves the relevant activation tensors, weight tensors, and corresponding gradient tensors, the tensors used during the training process of the DNN model only account for a small part of all tensors (less than 10%, about 1% on average); and since a tensor in the DNN model is usually only accessed once during the forward propagation and backward propagation respectively, the interval between tensor accesses in the DNN model is relatively long (for example, in the CNN model, the inactive period of more than 60% of the tensors exceeds 107 microseconds, and in the Transformer model, about 50% of the tensors have an inactive period greater than 105 microseconds).

[0160] In the DNN model, the size distribution and inactive period duration distribution of inactive tensors have a large span. For example, in the Inceptionv3 - 512 model, the size of inactive tensors ranges from 10KB to 2.7GB, and the inactive period duration ranges from 10 microseconds to 100 seconds, with a large span. Based on the tensor offloading rules, tensors with greater tensor offloading benefits can be offloaded, thereby significantly reducing the occupancy of video memory resources.

[0161] Figure 13 An exemplary flowchart of the tensor management method based on the present disclosure is shown.

[0162] Figure 13 In the shown flowchart, taking the simultaneous training of three tasks as an example, the tensor management method may include the following content.

[0163] Step S410, in response to receiving three model training tasks, add these three model training tasks to the waiting queue.

[0164] In some embodiments of the present disclosure, the first model training task among the three model training tasks in step S410 may be a ResNet50 image classification model training task; the second model training task may be a BERT Base text classification model training task; the third model training task may be a UNet medical image segmentation model training task. Through step S410, the above three model training tasks are initially in the waiting queue.

[0165] Step S420, obtain the real - time remaining amount of video memory and the video memory required for each model training task in the first queue.

[0166] In the embodiments of the present disclosure, the real - time remaining amount of video memory obtained through step S410 is 16GB, the video memory required for the ResNet50 image classification model training task is about 6GB, the video memory required for the BERT Base text classification model training task is about 8GB, and the video memory required for the UNet medical image segmentation model training task is about 10GB.

[0167] Step S430: Determine the to-be-trained tasks and delayed training tasks from the model training tasks in the first queue according to the real-time remaining amount and the video memory required by each model training task in the first queue.

[0168] In some embodiments of the present disclosure, the processing rule of the waiting queue can be to preferentially put the tasks with smaller required video memory into the execution queue, etc. Since the sum of the video memory required by the ResNet50 image classification model training task and the BERT Base text classification model training task (about 14GB) is less than the real-time remaining amount, the ResNet50 image classification model training task and the BERT Base text classification model training task can be used as the to-be-trained tasks, and the UNet medical image segmentation model training task can be used as the delayed training task.

[0169] Step S440: Migrate the delayed training tasks from the waiting queue to the delay queue, and migrate the to-be-trained tasks from the waiting queue to the execution queue.

[0170] In some embodiments of the present disclosure, after migrating the delayed training tasks from the waiting queue to the delay queue through step S440, the delayed training tasks can be migrated back to the waiting queue when the real-time remaining amount of the video memory meets the processing requirements of the UNet medical image segmentation model training task, etc.

[0171] Step S450: Control the execution of the model training tasks in the execution queue.

[0172] In some embodiments of the present disclosure, step S450 can control the execution of the ResNet50 image classification model training task and the BERT Base text classification model training task according to the priorities of the model training tasks in the execution queue.

[0173] Step S460: Monitor whether there is a risk of video memory over-commitment.

[0174] In some embodiments of the present disclosure, assume that the ResNet50 image classification model training task is currently being executed, the BERT Base text classification model training task is to be executed, and the real-time remaining amount of the current video memory is 3GB. Therefore, it is monitored through step S460 that there is a risk of video memory over-commitment.

[0175] Step S470: Obtain the task migration benefit ratio of each model training task in the execution queue.

[0176] Step S480: Determine the tasks to be migrated out from the model training tasks in the execution queue according to the task migration benefit ratio.

[0177] Step S490: Migrate the tasks to be migrated out from the execution queue to the pause queue.

[0178] In some embodiments of the present disclosure, since there is only one task left unexecuted in the execution queue in this example, steps S470 to S480 can be skipped, and the BERT Base text classification model training task can be directly migrated from the execution queue to the pause queue through step S490.

[0179] The process of task scheduling between different queues through steps S410 to S490 can be as Figure 14 shown.

[0180] Step S500, during the execution of the ResNet50 image classification model training task, obtain the training progress and tensor offloading rules of the model.

[0181] Step S510, obtain the stage offloading rules of the training progress from the tensor offloading rules.

[0182] Step S520, obtain the resident tensors of the target model in the video memory.

[0183] Step S530, determine the tensors to be offloaded from the resident tensors according to the stage offloading rules.

[0184] Step S540, offload the tensors to be offloaded to the target storage device.

[0185] For the tensor management method provided by the present disclosure, the model training tasks in the execution queue can determine the execution order according to the preset priority rules; when it is monitored that there is a risk of video memory over - score, the tasks in the execution queue are added to the pause queue to release a part of the video memory space to avoid excessive use of the video memory, thereby preventing the impact on training due to video memory over - score.

[0186] The present disclosure also provides a tensor management device (corresponding to the tensor management method). Figure 15 The schematic diagram showing the hardware implementation manner using a processing system is shown.

[0187] As Figure 16As shown, the hardware structure of the electronic device / apparatus can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and overall design constraints. Bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. Bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one connecting line is shown in this figure, but it does not mean there is only one bus or one type of bus.

[0188] For ease of description, certain steps of the above method are described corresponding to modules. It should be understood that the corresponding modules for executing one or several steps of the above method can be one or more hardware modules specifically configured to execute the corresponding steps, or implemented by a processor configured to execute the corresponding steps, or stored in a computer-readable medium for implementation by a processor, or implemented through a certain combination.

[0189] As Figure 15 shown, the tensor management device includes a progress acquisition module 1010, a rule acquisition module 1020, a residency acquisition module 1030, an offloading acquisition module 1040, and a tensor offloading module 1050.

[0190] The progress acquisition module 1010 is used to acquire the training progress of the target model and the tensor offloading rule during the training process of the target model; the tensor offloading rule is generated based on the offloading benefit-cost ratio determined according to the tensor life cycle of the target model.

[0191] The rule acquisition module 1020 is used to acquire the stage offloading rule of the training progress from the tensor offloading rule.

[0192] The residency acquisition module 1030 is used to acquire the resident tensors of the target model in the video memory.

[0193] The offloading acquisition module 1040 is used to determine the tensors to be offloaded from the resident tensors according to the stage offloading rule.

[0194] The tensor offloading module 1050 offloads the tensor to be offloaded to the target storage device. The specific implementation of each module in the above device may refer to the implementation process of the corresponding steps in the above method embodiments of the present disclosure, which will not be elaborated here.

[0195] The present disclosure also provides a readable storage medium storing a computer program which, when executed by a processor, is used to implement the above method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0196] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.

[0197] The computer program or instructions can be stored in a readable storage medium or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0198] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0199] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0200] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0201] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0202] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0203] Those skilled in the art should understand that the above embodiments are merely for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or variations can be made based on the above disclosure, and these changes or variations are still within the scope of the present disclosure.

Claims

1. A tensor management method, characterized in that, include: During the target model training process, the training progress and tensor unloading rules of the target model are obtained; the tensor unloading rules are generated according to the unloading benefit-cost ratio determined according to the tensor life cycle of the target model; Obtaining a stage unloading rule of the training progress from the tensor unloading rule; Get the resident tensor of the target model in the video memory; Determining a to-be-unloaded tensor from the resident tensors according to the stage unloading rule; as well as Unload the tensor to be unloaded to a target storage device.

2. The tensor management method according to claim 1, characterized in that Before training the target model, it also includes: Obtaining tensors of the target model and various stages of the training process of the target model; Respectively obtain the tensor lifecycle of each tensor of the target model; Determine the offloading benefit-cost ratio of each tensor at each stage according to the tensor life cycle of each tensor; and Generate tensor offloading rules based on the offloading benefit-cost ratio of each tensor in each stage.

3. The tensor management method according to claim 2, characterized in that For any tensor of the target model, determining the unloading benefit-cost ratio of the tensor at each stage according to the tensor life cycle of each tensor includes: Get the tensor size and tensor unloading cost of the tensor; Obtain the tensor idle time of the tensor at each stage according to the tensor life cycle of the tensor; Determine the tensor unloading benefit of the tensor at each stage according to the product of the tensor size and the tensor idle time of the tensor at each stage; and The unloading benefit-overhead ratio of the tensor at each stage is determined according to the quotient of the tensor unloading benefit of the tensor at each stage and the tensor unloading overhead.

4. The tensor management method according to claim 3, wherein The obtaining of the tensor size and tensor unloading overhead of the tensor includes: Get the tensor size and network bandwidth of the tensor; and The tensor offloading overhead of the tensor is determined according to a quotient of twice the tensor size of the tensor and the network bandwidth.

5. The tensor management method according to any one of claims 1 to 3, characterized in that After unloading the to-be-unloaded tensor to the target storage device, the method further includes: Obtain the tensor lifecycle and current time of the tensor to be unloaded; Determine the remaining time according to the tensor life cycle of the tensor to be unloaded and the current time; Determine whether the remaining time is greater than a time threshold; and In response to the remaining time of a target tensor in the to-be-unloaded tensors being not greater than the time threshold, migrating the target tensor to the video memory.

6. The tensor management method according to claim 5, wherein The step of migrating the target tensor to the video memory includes: Adding the target tensor to a prefetch queue; and Asynchronously fetching the tensor in the pre-fetch queue and adding it to the video memory.

7. The tensor management method according to any one of claims 1 to 3, characterized in that Before training the target model, it also includes: In response to receiving the model training task, adding the model training task to the first queue; Obtaining the real-time remaining amount of the video memory and the video memory required by each model training task in the first queue; Determine, from the model training tasks in the first queue, tasks to be trained and delayed training tasks according to the real-time remaining amount and the video memory required by each model training task in the first queue; Migrating the delayed training task from the first queue to the second queue, and migrating the task to be trained from the first queue to the third queue; the second queue is used to manage the delayed training task, and migrate the delayed training task back to the first queue when the first migration indication is triggered; and Control the execution of the tasks to be trained in the third queue, where the tasks to be trained in the third queue include the tasks for training the target model.

8. An electronic device, characterized in that, Comprising: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, such that the processor executes the tensor management method according to any one of claims 1 to 7.

9. A readable storage medium, characterized in that, Execution instructions are stored in the readable storage medium, and when the execution instructions are executed by a processor, they are used to implement the tensor management method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the tensor management method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Memory management method and device, electronic equipment and computer readable storage medium

    CN112559165A

  • Unloading training optimization method and device, electronic equipment and storage medium

    CN118796514A

  • Method for model training, host and storage device

    CN119849583A

  • Method and system for distributed data management

    WO2021010896A1

  • Memory optimization method and apparatus used for neural network compilation

    WO2024065867A1

Cited By

  • Large model reasoning tensor unloading method and system in resource limited scene

    CN120540744A

  • Neural network intermediate data unloading method and system based on GDS technology

    CN122088579A

  • A method and system for offloading neural network intermediate data based on GDS technology

    CN122088579B