Residual video memory optimization method and system for distributed training

By dividing the DNN model into reusable groups and specific layers, and utilizing the similarity of adjacent residual layers for average residual replacement and dimensionality compression, the problem of limited memory optimization effect in distributed training is solved, and the memory overhead is reduced and the model training efficiency is improved.

CN120671746APending Publication Date: 2025-09-19HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510826366.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have limited graphics memory optimization effects in distributed deep neural network training, resulting in large graphics memory overhead and affecting training efficiency.

Method used

The DNN model is divided into reusable groups and specific layers. The similarity of adjacent residual layers is used to perform average residual replacement, and the dimension of specific layers is compressed. Combined with existing memory optimization technology, the memory overhead is further reduced.

Benefits of technology

Through inter-layer and intra-layer memory optimization, memory overhead is reduced, while ensuring the convergence performance and accuracy of model training and improving training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671746A_ABST
    Figure CN120671746A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer storage, and particularly discloses a distributed training-oriented residual video memory optimization method and system, and the method comprises the steps: obtaining a difference value between an original gradient and a compression gradient in the training of a distributed deep neural network DNN model of compression communication as a residual; the DNN model is divided into a reusable group and a specific layer, the reusable group comprises a plurality of adjacent residual error layers with the same structure, and the specific layer is a layer not included in the reusable group; based on the average residual error in the reusable group, reducing the residual error in the reusable group, and performing dimension compression on the residual error of the specific layer; and updating parameters of the DNN model based on the reduced residual error and the residual error after dimension compression. According to the method, the video memory overhead of distributed training can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer storage technology, and more specifically, relates to a residual video memory optimization method and system for distributed training. Background Art

[0002] In order to reduce the memory overhead of distributed deep neural network (DNN) training, various methods have been proposed in the prior art. For example, the main idea of ​​activation recomputation is to not store most activation values ​​and recalculate them when necessary to calculate the gradient during the backward pass. However, this approach introduces a large time overhead and has a significant impact on training efficiency. ZeRO-DP partitions the model state instead of copying it to reduce the memory redundancy of the model state. StrongHold expands the maximum trainable model size by dynamically offloading data from the GPU to the CPU, and can be used to fine-tune pre-trained large models with limited GPU resources. However, the above methods have limited effect in reducing memory overhead. Summary of the Invention

[0003] In response to the shortcomings of the existing technology, the purpose of this application is to provide a residual video memory optimization method and system for distributed training, aiming to solve the problem that the existing video memory optimization methods have limited effect in reducing video memory overhead.

[0004] To achieve the above objectives, in a first aspect, the present application provides a residual video memory optimization method for distributed training, comprising: Obtain the difference between the original gradient and the compressed gradient in the distributed deep neural network (DNN) model training of compressed communication as the residual; Divide the DNN model into a reusable group and specific layers, wherein the reusable group includes multiple adjacent residual layers with the same structure, and the specific layers are layers not included in the reusable group; Based on the average residual in the reusable group, reducing the residuals in the reusable group and performing dimension compression on the residuals of the specific layer; Based on the reduced residual and the dimensionally compressed residual, the parameters of the DNN model are updated.

[0005] This application divides adjacent residual layers of the same type into reusable groups, and utilizes the similarity of adjacent residual layers of the same type to replace the original residuals with the average residuals in the reusable groups, thereby reducing the memory overhead of inter-layer residuals and ensuring the convergence performance of model training. It also further reduces the memory occupancy of intra-layer residuals by dimensionality compression of non-reusable specific layer residuals, which can be combined with existing memory optimization technology to further reduce memory overhead.

[0006] According to a residual memory optimization method for distributed training provided by the present application, the residual in the reusable group is reduced based on the average residual in the reusable group, including: replacing all residuals in the reusable group with the average residual; Reduction of residuals in reusable groups with replacement.

[0007] Based on the observation that adjacent residual layers of the same type have similar probability distributions, this application stores only the average value of the residuals in the reusable group instead of the original residuals, reducing the memory overhead of the inter-layer residuals; at the same time, by utilizing the similarity of adjacent residual layers of the same type, the average value residual ensures the convergence performance of the model training.

[0008] According to a residual video memory optimization method for distributed training provided by this application, the method also includes: Calculating the weight of each residual in the reusable group based on the importance of different residual layers in the reusable group; A weighted average residual is calculated based on the weights as the average residual.

[0009] This application calculates weights based on the importance of different residual layers in the reusable group and further calculates the weighted average residual, avoiding the decline in convergence performance and accuracy caused by the different importance of different residual layers in the reusable group, and ensuring the model accuracy while reducing the residual memory usage.

[0010] According to a residual memory optimization method for distributed training provided by this application, the dimensionality compression of the residual of the specific layer includes: Based on the dimension importance of the specific layer, the two-dimensional residual tensor is sparsely compressed.

[0011] This application selects the residual compression method of row and column indexes and overlapping elements according to the importance of dimensions, avoiding the problem of decreased model accuracy caused by missing dimension indexes, and ensuring that the model convergence and accuracy are not affected.

[0012] According to a residual memory optimization method for distributed training provided by the present application, the parameters of the DNN model are updated based on the reduced residual and the dimensionally compressed residual, including: Storing the reduced residual and the dimensionally compressed residual in local GPU memory, and compensating the original gradient based on the reduced residual and the dimensionally compressed residual; Compress the original gradient; Update the parameters of the DNN model based on the compressed original gradients.

[0013] In a second aspect, the present application provides a residual memory optimization system for distributed training, comprising: An acquisition module is used to obtain the difference between the original gradient and the compressed gradient in the training of the distributed deep neural network (DNN) model with compressed communication as the residual; A partitioning module is used to divide the DNN model into reusable groups and specific layers, wherein the reusable group includes multiple adjacent residual layers with the same structure, and the specific layer is a layer not included in the reusable group; a reduction module for reducing the residuals in the reusable group based on the average residual in the reusable group and performing dimension compression on the residuals of the specific layer; The update module is used to update the parameters of the DNN model based on the reduced residual and the dimensionally compressed residual.

[0014] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the residual video memory optimization method for distributed training described in the first aspect or any possible implementation of the first aspect.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the residual video memory optimization method for distributed training described in the first aspect or any possible implementation of the first aspect.

[0016] In a fifth aspect, the present application provides a computer program product, which, when running on a processor, enables the processor to execute the residual video memory optimization method for distributed training described in the first aspect or any possible implementation of the first aspect.

[0017] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0018] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies: This application divides adjacent residual layers of the same type into reusable groups, and utilizes the similarity of adjacent residual layers of the same type to replace the original residuals with the average residuals in the reusable groups, thereby reducing the memory overhead of inter-layer residuals and ensuring the convergence performance of model training. It also further reduces the memory occupancy of intra-layer residuals by dimensionality compression of non-reusable specific layer residuals, which can be combined with existing memory optimization technology to further reduce memory overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 Schematic diagram of the residual memory optimization method for distributed training provided in an embodiment of the present application; Figure 2 This is a schematic diagram of reusing the average value of inter-layer residual errors provided by an embodiment of the present application; Figure 3 Schematic diagram of intra-layer residual sparsification compression provided by an embodiment of the present application; Figure 4 is a schematic diagram of the error feedback process provided in an embodiment of the present application; Figure 5 This is a system architecture diagram provided by an embodiment of the present application; Figure 6 Schematic diagram of the structure of the residual memory optimization system for distributed training provided in an embodiment of the present application; Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0022] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.

[0023] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0024] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0025] First, let’s introduce the following contents: A DNN model essentially consists of multiple neural network layers, containing thousands or even millions of parameters. To achieve model convergence (i.e., the model training process reaches a stable state), DNN training iterates over a dataset over multiple cycles (epochs), with each epoch consisting of multiple iterations. During each iteration, the model performs forward and backward propagation to generate gradients for each layer, thereby updating the model parameters.

[0026] Existing DNN training systems typically distribute training tasks across multiple nodes (i.e., GPUs) and employ parallelization to accelerate training and handle large-scale datasets and models. In each iteration, the dataset is divided into multiple mini-batches and assigned to different nodes. Each node calculates gradients based on its assigned mini-batch and then synchronizes these gradients with other nodes to jointly update the global model parameters.

[0027] During DNN training, gradient communication between nodes often results in significant communication overhead, which severely limits DNN training performance. Therefore, many distributed DNN training systems adopt gradient compression to reduce the amount of data transmitted during gradient synchronization. Gradient compression methods can be roughly divided into two categories: sparsification, where a subset of gradient elements is selected to generate a sparse vector for communication, and quantization, where the number of bits of each gradient element is reduced to reduce the amount of communication. Top-k sparsification has been shown to be one of the simplest and most effective gradient compression methods. Top-k sparsification selects the k gradient elements with the largest absolute values ​​from all elements of the gradient, while setting the remaining elements to zero.

[0028] Due to the lossy nature of gradient compression, it often results in a decrease in model accuracy. Therefore, residuals are widely used in many gradient compression methods to ensure convergence and accuracy during model training. The process of compensating gradients using residuals is often referred to as error feedback. The gradients not propagated in iteration t are added as residuals to iteration t + 1, compensating for the loss in DNN model accuracy caused by gradient compression. Specifically, in each iteration, backpropagation calculates the original gradients; the residuals are accumulated with the original gradients to obtain the compensated gradients; the compensated gradients are compressed for communication and model updates; the difference between the compensated gradients and the compressed gradients is calculated, and the residual is obtained and stored in GPU memory.

[0029] However, residuals incur significant memory overhead, as each training iteration requires saving the residuals to each node's local GPU memory (i.e., video memory). Since gradient compensation essentially involves adding the original gradient g to the residual r, the memory usage of the residuals matches the size of the gradients. Furthermore, since gradients have a one-to-one correspondence with model parameters, the residuals are the same size as the gradients and model parameters, resulting in significant memory overhead.

[0030] Next, combine Figure 1-Figure 5 The residual video memory optimization method for distributed training provided in the embodiments of the present application is introduced.

[0031] Figure 1 This is a flow chart of the residual memory optimization method for distributed training provided by the embodiment of the present application, such as Figure 1 As shown, the method includes the following steps: Step 100: Obtain the difference between the original gradient and the compressed gradient in the distributed deep neural network (DNN) model training of compressed communication as the residual; Obtaining residuals in distributed DNN training with compressed communication. In each iteration, the difference between the original gradient and the compressed gradient of each node is calculated as the residual, which represents the gradient that was not transmitted by each node.

[0032] Step 110: Divide the DNN model into a reusable group and specific layers, wherein the reusable group includes multiple adjacent residual layers with the same structure, and the specific layer is a layer not included in the reusable group; According to the structural characteristics of the DNN model, it is divided into reusable groups containing multiple adjacent residual layers with the same structure, and specific layers that cannot be reused.

[0033] Step 120 , based on the average residual in the reusable group, reducing the residuals in the reusable group and performing dimension compression on the residuals of a specific layer; For reusable groups and specific layers, inter-layer and intra-layer residual reduction are performed respectively. For each reusable group, inter-layer residual reduction is performed. Specifically, the average value of all adjacent residual layers with the same structure in the group is calculated as the average residual, and the residuals in the reusable group are reduced based on the average residual; for specific layers, intra-layer residual reduction is performed. Specifically, the residuals of the specific layers are dimensionality compressed.

[0034] Step 130: Update the parameters of the DNN model based on the reduced residual and the dimensionally compressed residual.

[0035] The residual memory optimization method for distributed training provided in this application divides adjacent residual layers of the same type into reusable groups, utilizes the similarity of adjacent residual layers of the same type, replaces the original residuals with the average residuals in the reusable groups, reduces the memory overhead of inter-layer residuals, ensures the convergence performance of model training, and further reduces the memory occupancy of intra-layer residuals by dimensionality compression of non-reusable specific layer residuals. It can be combined with existing memory optimization technology to further reduce memory overhead.

[0036] In some embodiments, step 120 specifically includes: Step 1201, replacing all residuals in the reusable group with the average residual; Step 1202: Reduce the residual in the replaced reusable group.

[0037] According to different layer types, the DNN model is divided into m reusable groups, each group contains residual layers, where the reusable group number j satisfies 1≤j≤m, and the residual values ​​in all reusable groups are calculated , and through the jth group The mean of the residuals is calculated (average residual, referred to as ), and store it in GPU memory instead of storing A residual.

[0038] Figure 2 This is a schematic diagram of reusing the average value of inter-layer residual errors provided by the embodiment of the present application. Figure 2 As shown in the figure, there are 2 reusable groups, each group contains 3 adjacent neural network layers with the same structure, and the corresponding gradients are and In each iteration, two sets of corresponding residuals are calculated respectively. and Calculating the mean of the residuals in the two groups yields and , and store it in GPU memory instead of storing the original 6 residuals. and Compensate two groups of 6 original gradients respectively and .

[0039] In some embodiments, the method further comprises: Based on the importance of different residual layers in the reusable group, the weight of each residual in the reusable group is calculated; Calculate the weighted mean residual based on the weights as the average residual.

[0040] Preferably, the weighted average residual (weighted average residual, abbreviated as ), rather than a simple residual mean , where different weights represent the importance of different residual layers within the reusable group; According to the L1 and L2 norms, the weights of the residuals in each reusable group are calculated as follows:

[0041] in, Residual The weight of and Respectively represent the residual The L1 and L2 norms of the residual number i satisfy 1≤i≤ ; Residual The number of elements of Divide by to be of the same order of magnitude; Indicates the importance coefficient, the default value is 0.5.

[0042] According to the calculated residual weights, the weighted average residual of each reusable group is calculated. The formula is:

[0043] in, represents the weighted mean residual of the j-th reusable group.

[0044] In some embodiments, step 120 specifically includes: Step 1203: Based on the dimensional importance of a specific layer, the two-dimensional residual tensor is sparsely compressed.

[0045] For specific layers in the DNN model, such as the fully connected layer and the word embedding layer, the two-dimensional residual tensor is sparsely compressed according to the importance of the dimension. The residual tensor of , select the most important ones according to the L1 norm and Only the selected row and column indices and the corresponding overlapping elements are stored, thereby achieving intra-layer two-dimensional residual row and column sparse compression.

[0046] Figure 3 This is a schematic diagram of intra-layer residual sparse compression provided by an embodiment of the present application, such as Figure 3 As shown, Figure 3 (a) represents an 8×6 residual tensor of a specific layer, with 8 rows and 6 columns; Figure 3(b) shows the compression result of 25% compression rate from one dimension, selecting R0 and R6 and their corresponding residual elements from 8 rows; Figure 3 (c) represents a two-dimensional sparse compression method, which selects four rows R0, R3, R4, and R6 and three columns C1, C4, and C5, and saves the 12 overlapping elements in the selected rows and columns, which also achieves the same Figure 3 (b) The same 25% compression rate saves more dimensional information, which can improve the model accuracy.

[0047] In some embodiments, step 140 specifically includes: Step 1401: Store the reduced residual and the dimensionally compressed residual in the local GPU memory, and compensate the original gradient based on the reduced residual and the dimensionally compressed residual. Step 1402, compressing the original gradient; Step 1403: Update the parameters of the DNN model based on the compressed original gradients.

[0048] For the jth reusable group, reuse To compensate for all Original gradient Optionally, after compensating the original gradient, gradient compression and synchronous communication are performed, and the parameters of the DNN model are updated.

[0049] Optionally, synchronous communication of gradients is implemented through collective communication primitives, such as AllGather, and parameter updates are performed using the following formula:

[0050] in, and are the neural network parameters and gradients of the i-th working node at the t-th iteration, is the learning rate, is the number of working nodes, represents gradient compression (e.g. Top-k), Represents gradient decompression, reconstructing the shape and size of the original gradient.

[0051] Figure 4 is a schematic diagram of the error feedback process provided by the embodiment of the present application, such as Figure 4 As shown, in each iteration, the error feedback operation is as follows: 1a. Back propagation calculation to obtain the original gradient ; 2a. Residual Accumulate with the original gradient to get the compensated gradient , the specific formula is:

[0052] in, and are the residual and compensation gradient of the i-th working node at the t-th iteration respectively; 3a. Compress the gradient to obtain , and communicate; 4a. Update model parameters using aggregated gradients. 5a. Update the residual and calculate the difference between the compensation gradient and the compression gradient. The formula is as follows:

[0053] Figure 5 This is a system architecture diagram provided by the embodiment of the present application, such as Figure 5 As shown, in one embodiment of the present application, the residual video memory optimization method for distributed training provided by the present application includes the following steps: 1b. Gradient Compression: This section implements three classic gradient compression methods: Top-k sparsification, DGC sparsification, and QSGD quantization. It also provides a gradient compression API, including a compression function (compress), which sparsifies or quantizes the gradients in each iteration while preserving the number of elements and shape of the original gradients; and a decompression function (decompress), which reconstructs the compressed gradients into the original gradients for parameter updates.

[0054] 2b. Residual reduction: Two residual reduction methods are used: inter-layer residual reuse and intra-layer residual compression. Inter-layer residual reuse reuses the weighted average of the residuals in reusable groups of adjacent neural network layers of the same shape instead of storing all the residuals in the group. Intra-layer residual compression performs sparse compression on non-reusable two-dimensional residual tensors along two dimensions, storing only the important dimension indices and overlapping elements.

[0055] 3b. Error feedback. The error feedback API is implemented, including a residual update function (update), which calculates the difference between the gradient and the compressed gradient as the residual in each iteration and stores it in GPU memory; and a gradient compensation function (compensate), which accumulates the residual stored in local GPU memory with the original gradient to compensate the original gradient.

[0056] 4b. Gradient communication. Three collective communication primitive APIs are provided, namely Allreduce, Allgather, and Allgather_Faster, to achieve efficient communication and synchronous aggregation of compressed gradients.

[0057] Figure 6 This is a schematic diagram of the structure of a residual memory optimization system for distributed training provided by an embodiment of the present application. Figure 6 As shown, the system includes an acquisition module 610, a division module 620, a reduction module 630 and an update module 640, wherein: An acquisition module 610 is configured to obtain a difference between an original gradient and a compressed gradient in training of a distributed deep neural network (DNN) model for compressed communication as a residual; A partitioning module 620 is configured to partition the DNN model into reusable groups and specific layers, wherein the reusable groups include multiple adjacent residual layers with the same structure, and the specific layers are layers not included in the reusable groups; A reduction module 630 for reducing the residuals in the reusable group based on the average residual in the reusable group and performing dimensionality compression on the residuals of a specific layer; The updating module 640 is used to update the parameters of the DNN model based on the reduced residual and the dimensionally compressed residual.

[0058] Based on the method in the above embodiment, Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7 As shown, an embodiment of the present application provides an electronic device, which may include: a processor (processor) 710, a communication interface (Communications Interface) 720, a memory (memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logic instructions in the memory 730 to execute the residual video memory optimization method for distributed training in the above embodiment.

[0059] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the residual video memory optimization method for distributed training described in various embodiments of the present application.

[0060] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the residual video memory optimization method for distributed training in the above embodiment.

[0061] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the residual video memory optimization method for distributed training in the above embodiment.

[0062] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0063] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.

[0064] The above embodiments can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0065] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0066] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A residual memory optimization method for distributed training, characterized in that: include: Obtain the difference between the original gradient and the compressed gradient in the distributed deep neural network (DNN) model training of compressed communication as the residual; Divide the DNN model into a reusable group and specific layers, wherein the reusable group includes multiple adjacent residual layers with the same structure, and the specific layer is a layer not included in the reusable group; Based on the average residual in the reusable group, reducing the residuals in the reusable group and performing dimension compression on the residuals of the specific layer; Based on the reduced residual and the dimensionally compressed residual, the parameters of the DNN model are updated.

2. The residual video memory optimization method for distributed training according to claim 1, characterized in that: The reducing the residuals in the reusable group based on the average residuals in the reusable group includes: replacing all residuals in the reusable group with the average residual; Reduction of residuals in reusable groups with replacement.

3. The residual video memory optimization method for distributed training according to claim 1, characterized in that: The method further comprises: Calculating the weight of each residual in the reusable group based on the importance of different residual layers in the reusable group; A weighted average residual is calculated based on the weights as the average residual.

4. The residual video memory optimization method for distributed training according to claim 1, characterized in that: The performing dimension compression on the residual of the specific layer includes: Based on the dimension importance of the specific layer, the two-dimensional residual tensor is sparsely compressed.

5. The residual video memory optimization method for distributed training according to claim 1, characterized in that: The updating of the parameters of the DNN model based on the reduced residual and the dimensionally compressed residual includes: Storing the reduced residual and the dimensionally compressed residual in local GPU memory, and compensating the original gradient based on the reduced residual and the dimensionally compressed residual; Compress the original gradient; Update the parameters of the DNN model based on the compressed original gradients.

6. A residual memory optimization system for distributed training, characterized in that: include: An acquisition module is used to obtain the difference between the original gradient and the compressed gradient in the training of the distributed deep neural network (DNN) model with compressed communication as the residual; A partitioning module is used to divide the DNN model into reusable groups and specific layers, wherein the reusable group includes multiple adjacent residual layers with the same structure, and the specific layer is a layer not included in the reusable group; a reduction module for reducing the residuals in the reusable group based on the average residual in the reusable group and performing dimension compression on the residuals of the specific layer; The update module is used to update the parameters of the DNN model based on the reduced residual and the dimensionally compressed residual.

7. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the residual video memory optimization method for distributed training as described in any one of claims 1-5.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor executes the residual video memory optimization method for distributed training as described in any one of claims 1 to 5.

9. A computer program product, characterized in that When the computer program product runs on a processor, the processor executes the residual video memory optimization method for distributed training as described in any one of claims 1 to 5.