Memory Management Method, Device, and Storage Medium
By analyzing the computational subgraph of the neural network model, determining the execution device of the operator and performing memory optimization management, the problem of low memory application and data copying efficiency of AI tasks on different hardware platforms is solved, and the model operation performance is improved.
Patent Information
- Application Number
- CN202311774780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-12-21
AI Technical Summary
When existing AI tasks run on different hardware platforms, they cannot effectively obtain the operation of model operators on GPU or CPU devices, resulting in large overhead of memory applications and data copying between devices, affecting model operation performance.
By analyzing the intermediate expression corresponding to the computational subgraph of the neural network model, the execution devices of each operator are determined, and memory optimization management is carried out, including memory pre-application, data copying and optimization processing, and the target executable file is generated to improve the efficiency of data copying between memory applications and devices.
It improves the efficiency of memory application and data copying between devices, optimizes memory space, reduces data transmission overhead, and improves the operating performance of the model on the execution device.
Smart Images

Figure CN117667424B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of artificial intelligence (AI), and in particular, to a memory management method, apparatus, and storage medium. Background Art
[0002] With the continuous development of AI technology, the AI compute-intensive scenarios have put forward higher requirements for computing power, and hardware manufacturers have proposed different hardware solutions to provide stronger computing power. When AI runs and accelerates on different hardware platforms, corresponding backend adaptation needs to be performed according to different hardware platforms, and the AI compiler ensures the accelerated operation of AI tasks on specific hardware platforms. In order to make AI tasks run more efficiently on the hardware platform, a specific solution design needs to be carried out for the memory management method in the AI compiler for a specific hardware platform.
[0003] In the existing solutions, usually the running situation of model operators on specific devices such as GPUs or CPUs cannot be obtained during the compilation process, so it is assumed that all the computing operators in the entire model run on one-sided devices. If there are operators in the model that need to run on other-sided devices, the memory of the corresponding devices needs to be repeatedly allocated, and the memory data copy operations between multiple-sided devices need to be performed multiple times. The overhead of memory allocation and data copy between devices is particularly large, which greatly affects the performance of model operation. Therefore, there is an urgent need for a new memory management solution to improve the efficiency of memory allocation and data copy between devices and improve the performance of the model running on the device. Summary of the Invention
[0004] In view of this, the present disclosure provides a memory management method, apparatus, and storage medium.
[0005] According to one aspect of the present disclosure, a memory management method is provided. The method can be used for an analysis device, and the method includes:
[0006] Based on the intermediate representation corresponding to the computation subgraph of the neural network model, determine the execution device for each operator in the intermediate representation respectively;
[0007] Based on the execution devices corresponding to the operators in the intermediate representation, perform memory optimization management for different execution devices;
[0008] Based on the result of the memory optimization management, generate a target executable file on the execution device, and the target executable file is used for the execution device to perform training or inference based on the neural network model.
[0009] In a possible implementation manner, based on the intermediate representation corresponding to the computation subgraph of the neural network model, determining the execution device for each operator in the intermediate representation respectively includes:
[0010] Determine the execution device for executing the operator according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation.
[0011] In a possible implementation, the execution device includes a graphics processing unit (GPU) and a central processing unit (CPU). Determining the execution device for executing the operator according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation includes:
[0012] When the memory type of the input tensor and / or output tensor of the operator is GPU memory, determine that the execution device for executing the operator is the GPU; otherwise,
[0013] When the memory type of the input tensor and output tensor of the operator is CPU memory, determine that the execution device for executing the operator is the CPU.
[0014] In a possible implementation, the method further includes:
[0015] When there is a need to copy discontinuous data between the CPU and the GPU, or between GPUs, use a predetermined data parallel processing function in the GPU to copy the discontinuous data, where the discontinuous data is data that is not continuously stored in memory units.
[0016] In a possible implementation, based on the execution devices corresponding to the operators in the intermediate representation, perform memory optimization management for different execution devices, including:
[0017] Based on the execution devices corresponding to the operators in the intermediate representation, perform pre-application of memory under the corresponding execution devices.
[0018] In a possible implementation, based on the execution devices corresponding to the operators in the intermediate representation, perform memory optimization management for different execution devices, including:
[0019] For a predetermined first operator in the intermediate representation, cause the first operator to apply for memory in the corresponding execution device outside the scope of action. The first operator is an operator within a loop / nested operation, and the scope of action is the scope of the loop / nested operation.
[0020] In a possible implementation, based on the execution devices corresponding to the operators in the intermediate representation, perform memory optimization management for different execution devices, including:
[0021] For a predetermined second operator in the intermediate representation, based on the constraint conditions corresponding to the second operator, copy the input tensor and / or output tensor of the second operator between execution devices. The constraint conditions include that the memory type of the input tensor and / or output tensor of the second operator is memory on a predetermined execution device.
[0022] In one possible implementation, based on the execution devices corresponding to the operators in the intermediate representation, memory optimization management is performed for different execution devices, including:
[0023] For the read-write dependency relationships between the input tensors and output tensors among the operators in the intermediate representation, the input tensors and / or output tensors with read-write dependency relationships share the same piece of memory on the execution device.
[0024] In one possible implementation, the method further includes:
[0025] Perform optimization processing on the intermediate representation to determine the optimized intermediate representation. The optimization processing includes one or more of constant folding, redundancy elimination, and operator fusion;
[0026] Based on the optimized intermediate representation, determine the execution devices for executing the operators respectively for each operator in the optimized intermediate representation;
[0027] Based on the execution devices corresponding to the operators in the intermediate representation, memory optimization management is performed for different execution devices, including:
[0028] Based on the execution devices corresponding to the operators in the optimized intermediate representation, memory optimization management is performed for different execution devices.
[0029] In one possible implementation, the method can be used in the Just-In-Time (JIT) compilation stage under the PyTorch framework, and the intermediate representation is the Multi-Level Intermediate Representation (MLIR).
[0030] In one possible implementation, the computational subgraph is obtained by splitting the computational graph of the neural network model in the JIT stage, and the intermediate representation is obtained by transforming and expressing the operators of the computational subgraph.
[0031] According to another aspect of the present disclosure, a memory management device is provided. The device is used for an analysis device, and the device includes:
[0032] A first determination module, configured to determine the execution devices for executing the operators respectively for each operator in the intermediate representation based on the intermediate representation corresponding to the computational subgraph of the neural network model;
[0033] An optimization management module, configured to perform memory optimization management for different execution devices based on the execution devices corresponding to the operators in the intermediate representation;
[0034] A second determination module, configured to generate a target executable file on the execution device based on the result of the memory optimization management, and the target executable file is used for the execution device to perform training or inference based on the neural network model.
[0035] In a possible implementation, a first determination module is configured to:
[0036] Determine an execution device for executing an operator according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation.
[0037] In a possible implementation, the execution device includes a graphics processing unit (GPU) and a central processing unit (CPU). Determining an execution device for executing an operator according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation includes:
[0038] When the memory type of the input tensor and / or output tensor of the operator is GPU memory, determine that the execution device for executing the operator is the GPU; otherwise,
[0039] When the memory types of the input tensor and output tensor of the operator are CPU memory, determine that the execution device for executing the operator is the CPU.
[0040] In a possible implementation, the apparatus further includes:
[0041] A copy module is configured to, when there is a need to copy discontinuous data between the CPU and the GPU, or between GPUs, use a predetermined data parallel processing function in the GPU to copy the discontinuous data, where the discontinuous data is data that is not continuously stored in memory units.
[0042] In a possible implementation, an optimization management module is configured to:
[0043] Based on the execution devices corresponding to the operators in the intermediate representation, perform memory pre-allocation under the corresponding execution devices.
[0044] In a possible implementation, an optimization management module is configured to:
[0045] For a predetermined first operator in the intermediate representation, cause the first operator to apply for memory in the corresponding execution device outside the scope of action, where the first operator is an operator within a loop / nested operation, and the scope of action is the scope of the loop / nested operation.
[0046] In a possible implementation, an optimization management module is configured to:
[0047] For a predetermined second operator in the intermediate representation, based on the constraint conditions corresponding to the second operator, copy the input tensor and / or output tensor of the second operator between execution devices, where the constraint conditions include that the memory type of the input tensor and / or output tensor of the second operator is memory on a predetermined execution device.
[0048] In a possible implementation, an optimization management module is configured to:
[0049] Regarding the read-write dependency relationships between input tensors and output tensors of each operator in the intermediate representation, the input tensors and / or output tensors with read-write dependency relationships are reused in the same memory block on the execution device.
[0050] In a possible implementation, the apparatus further includes:
[0051] An optimization processing module, configured to perform optimization processing on the intermediate representation to determine the optimized intermediate representation, and the optimization processing includes one or more of constant folding, redundancy elimination, and operator fusion;
[0052] A third determination module, configured to determine the execution device for executing the operator for each operator in the optimized intermediate representation based on the optimized intermediate representation;
[0053] An optimization management module, configured to:
[0054] Perform memory optimization management for different execution devices based on the execution devices corresponding to the operators in the optimized intermediate representation.
[0055] In a possible implementation, the apparatus can be used in the Just-In-Time (JIT) compilation stage under the PyTorch framework, and the intermediate representation is the Multi-Level Intermediate Representation (MLIR).
[0056] In a possible implementation, the computational subgraph is obtained by splitting the computational graph of the neural network model in the JIT stage, and the intermediate representation is obtained by converting and expressing the operators of the computational subgraph.
[0057] According to another aspect of the present disclosure, there is provided a memory management apparatus, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory.
[0058] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0059] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above method.
[0060] According to the embodiments of the present application, by determining the execution device for each operator based on the intermediate representation corresponding to the computational subgraph on the analysis device, corresponding memory optimization management can be performed for different execution devices corresponding to each operator, thereby improving the efficiency of memory application and data copying between execution devices, optimizing the memory space, and reducing the overhead of data transmission. By generating a target executable file on the execution device based on the result of the memory optimization management to perform training or inference based on the neural network model on the execution device, the performance of the model running on the execution device can be improved.
[0061] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.
[0063] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present application.
[0064] Figure 2 A flowchart showing a memory management method according to an embodiment of the present application.
[0065] Figure 3 A flowchart showing a memory management method according to an embodiment of the present application.
[0066] Figure 4 A flowchart showing the process of performing memory optimization management according to an embodiment of the present application.
[0067] Figure 5 A schematic diagram showing the overall process of a memory management method according to an embodiment of the present application.
[0068] Figure 6 A structural diagram showing a memory management device according to an embodiment of the present application.
[0069] Figure 7 A block diagram showing an apparatus 1900 for memory management according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0071] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior or better than other embodiments.
[0072] In addition, for a better illustration of the present disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art should understand that the present disclosure may be practiced without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0073] With the continuous development of artificial intelligence (AI) technology, the AI compute-intensive scenarios have put forward higher requirements for computing power, and hardware manufacturers have proposed different hardware solutions to provide stronger computing power. For AI to run and be accelerated on different hardware platforms, corresponding backend adaptation needs to be performed according to different hardware platforms, and the AI compiler ensures the accelerated operation of AI tasks on specific hardware platforms. To make AI tasks run more efficiently on the hardware platform, a specific scheme design needs to be carried out for the memory management method in the AI compiler for a specific hardware platform. In existing schemes, the running situation of model operators on specific devices such as GPUs or CPUs cannot usually be obtained during the compilation process, so it is assumed that all computing operators in the entire model run on a single device. If there are operators in the model that need to run on other devices, the memory of the corresponding device needs to be repeatedly applied for, and memory data copying operations between multiple devices need to be performed multiple times. The overhead of memory application and data copying between devices is particularly large, thus greatly affecting the performance of model operation. Therefore, there is an urgent need for a new memory management scheme to improve the efficiency of memory application and data copying between devices and improve the performance of the model running on the device.
[0074] In view of this, the present application proposes a memory management method. The method of the embodiments of the present application can be used for an analysis device. On the analysis device, based on the intermediate representation corresponding to the computation subgraph, the execution device for executing each operator can be determined respectively for each operator, and corresponding memory optimization management can be performed for different execution devices corresponding to each operator, so as to improve the efficiency of memory application and data copying between execution devices, optimize the memory space, and reduce the overhead of data transmission. By generating a target executable file on the execution device based on the result of memory optimization management, for training or inference based on the neural network model on the execution device, the performance of the model running on the execution device can be improved.
[0075] Figure 1 A schematic diagram showing an application scenario according to an embodiment of the present application. As Figure 1As shown, the memory management system of the embodiment of the present application can be deployed on an analysis device, and the analysis device can be a server.
[0076] In the application scenario of the embodiment of the present application, the memory management system can be applicable to the just-in-time (JIT) compilation stage under the deep learning framework PyTorch. Among them, the Dynamo module under the PyTorch framework can be used to generate the computational subgraph of the neural network model, and the memory management system can determine the corresponding intermediate representation based on the computational subgraph. This intermediate representation is a multi-level intermediate representation (MLIR), and MLIR can be used by the AI compiler for compilation processing. During the compilation process using the AI compiler, the memory management system can, according to the method of the embodiment of the present application, respectively determine the execution devices for executing each operator in the MLIR (the execution devices can be servers or terminal devices, such as Figure 1 as shown, the execution devices can include a graphics processing unit (GPU) and a central processing unit (CPU)) for memory optimization management to generate the target executable file on the execution device. The analysis device can send the target executable file to the execution device, so that the execution device can perform training or inference based on the neural network model based on the target executable file.
[0077] The terminal device involved in the embodiment of the present application can be any one or more of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), and a vehicle-mounted device. The embodiment of the present application does not impose special restrictions on the specific type of the terminal device, and it can have wired or wireless communication functions to achieve interaction with other devices.
[0078] The server involved in the embodiment of the present application can be located locally or in the cloud, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., and has a wireless communication function. Among them, the wireless communication function can be set on the chip (system) or other components or assemblies of the server. The wireless communication function can be realized, for example, through mobile communication technologies such as 2G / 3G / 4G / 5G, as well as Wi-Fi, Bluetooth, frequency modulation (FM), data radio, satellite communication, etc. It can also communicate through a wired connection to achieve interaction with other devices.
[0079] The following is through Figures 2 - 5, introduce the memory management method of the embodiments of the present application.
[0080] Figure 2 The flowchart showing the memory management method according to an embodiment of the present application is shown. This method can be used in the above analysis device, such as Figure 2 As shown, this method may include:
[0081] Step S201, based on the intermediate expression corresponding to the computational sub-graph of the neural network model, respectively determine the execution device for executing the operators for each operator in the intermediate expression.
[0082] As Figure 1 shown, the method of the embodiments of the present application can be used in the JIT stage under the PyTorch framework, and this intermediate expression can be MLIR. The computational sub-graph can be obtained by splitting the computational graph of the neural network model in the JIT stage, and the intermediate expression can be obtained by transforming and expressing the operators of the computational sub-graph. For example, the running neural network model can be split into sub-graphs through the dynamo module under the PyTorch framework to obtain the computational sub-graph. The operators in the computational sub-graph can be processed through MLIR torchdialect to obtain the intermediate expression, that is, the multi-level intermediate expression MLIR.
[0083] Since the neural network model usually involves heterogeneous computing tasks across devices, the above execution devices can include GPUs and CPUs (see Figure 1 ), so data copying and memory application between GPUs and CPUs are involved. To improve the efficiency of subsequent data copying and reduce duplicate memory application, dynamic selection can be first performed on each operator in the intermediate expression, and the operator lists running on GPUs and CPUs are sorted out for device partitioning between GPUs and CPUs. Under the PyTorch framework, the execution device can be partitioned through the MLIR operator device analysis PASS (PASS can represent the optimization processing steps in the AI compiler).
[0084] Among them, the execution device for executing the operator can be determined according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate expression.
[0085] The memory type of the input tensor and output tensor of the operator can indicate whether the corresponding tensor is stored in GPU memory or stored in CPU.
[0086] Among them, when the memory type of the input tensor and / or output tensor of the operator is GPU memory, the execution device for executing the operator can be determined as GPU; otherwise, when the memory type of the input tensor and output tensor of the operator is CPU memory, the execution device for executing the operator can be determined as CPU.
[0087] Device partitioning can also be performed based on the calculation type of the operator and the running device to determine the execution device for executing the operator.
[0088] In a possible implementation, the execution device can be partitioned based on the criterion that the operator preferentially runs on the GPU. As long as the memory type of one of the input tensors and output tensors of the operator is GPU memory, the execution device for executing the operator can be considered as the GPU.
[0089] After partitioning the execution device for the operator, data transfer may be involved. For example, if the execution device of a certain operator is the GPU and the memory type of the input tensor / output tensor of the operator is CPU memory, before subsequent memory optimization management, it can be determined to preprocess the tensor data through a predetermined data parallel processing function in the GPU, and copy the corresponding tensor from the CPU memory to the GPU memory, so as to ensure that the corresponding operator can be correctly executed. Among them, the CPU running function is called for the operator with the execution device being the CPU, and the GPU running function is called for the operator with the execution device being the GPU. Or in the case where the input tensor / output tensor of some operators involves copying between GPU memories, before subsequent memory optimization management, it can be determined to preprocess the tensor data through a predetermined data parallel processing function in the GPU, and copy the corresponding tensor between GPU memories. The preprocessing method can be referred to as follows.
[0090] Optionally, the method may further include:
[0091] In the case where there is a need to copy discontinuous data between the CPU and the GPU, or between the GPU and the GPU, use a predetermined data parallel processing function in the GPU to copy the discontinuous data.
[0092] The discontinuous data is data that is discontinuously stored in memory units and can be tensor data. The predetermined data parallel processing function is, for example, the relevant execution method of the GPU Kernal, so as to reduce the calls of the H2D or D2D functions of the GPU and make the data transfer more efficient.
[0093] After executing S201, the low-level intermediate representation IR in MLIR can also be optimized, and the execution device of the operators in the optimized IR can be repartitioned according to the execution device corresponding to the operator determined above, so as to improve the performance of model operation. Refer to Figure 3 , which shows the flowchart of the memory management method according to an embodiment of the present application. As Figure 3 shown, based on the method shown in Figure 2 , the method may further include steps S301 and S302 executed after step S201:
[0094] Step S301: Optimize the intermediate representation to determine the optimized intermediate representation.
[0095] Among them, the intermediate representation can be a low-level intermediate representation in MLIR. The optimization process includes one or more of constant folding, redundancy elimination, and operator fusion. The methods of constant folding, redundancy elimination, operator fusion, etc. can be implemented based on existing technologies.
[0096] Step S302: Based on the optimized intermediate representation, determine the execution devices for each operator in the optimized intermediate representation respectively.
[0097] Among them, most of the operators in the optimized IR are the same as those in the pre-optimized IR. The execution devices corresponding to the operators in the optimized IR can be determined based on the execution devices corresponding to the operators obtained in the above S201. For other operators in the optimized IR, the methods shown in S201 can be used to determine their corresponding execution devices one by one based on the memory types of the input tensors and output tensors. In this way, not only the complete information determined through step S201 before optimization is retained, but also partial optimization is carried out on this basis, improving the determination result of the execution devices.
[0098] It can continue to refer back to Figure 2 :
[0099] Step S202: Based on the execution devices corresponding to the operators in the intermediate representation, perform memory optimization management for different execution devices.
[0100] Among them, according to the characteristics of the low-level IR in some specific scenarios, memory optimization management can be carried out through memory management PASS for specific scenarios, thereby improving the performance of model operation. The intermediate representation can be the optimized intermediate representation. Refer to Figure 3 , this step S202 may include:
[0101] Step S303: Based on the execution devices corresponding to the operators in the optimized intermediate representation, perform memory optimization management for different execution devices.
[0102] As Figure 3 shown, after step S303, the steps shown in Figure 2 can be continued to execute.
[0103] An exemplary process of performing memory optimization management on the analysis device can be referred to Figure 4 , which shows a flowchart of performing memory optimization management according to an embodiment of the present application. As Figure 4 shown, this step S202 may include:
[0104] Step S401: Based on the execution devices corresponding to each operator in the intermediate representation, perform pre-allocation of memory under the corresponding execution devices.
[0105] Among them, the intermediate representation can be an intermediate representation involving memory application. For the operators executed by GPU (i.e., the execution device corresponding to the operator is GPU), the gpu.alloc method can be used to apply for GPU memory for the operator; for the operators executed by CPU (i.e., the execution device corresponding to the operator is CPU), the memerf.alloc method can be used to apply for CPU memory for the operator.
[0106] Step S402: For a predetermined first operator in the intermediate representation, enable the first operator to apply for memory in the corresponding execution device outside the scope of action.
[0107] Among them, the first operator can be an operator within a loop / nested operation (such as the scf::while operator), and the scope of action is the scope of the loop / nested operation, that is, the operator executes within this scope.
[0108] The applied memory can be lifted outside this scope, so that the memory applied outside this scope can be reused in scenarios such as loops. In this way, there is no need to apply for memory multiple times and dilute the memory, thus achieving the effect of improving the memory reuse efficiency and reducing the device performance overhead.
[0109] During the process of memory optimization management, for some operators with constraint conditions on the memory type of input / output tensors, special processing can be performed, as described below.
[0110] Step S403: For a predetermined second operator in the intermediate representation, based on the constraint conditions corresponding to the second operator, copy the input tensor and / or output tensor of the second operator between execution devices.
[0111] The operators in the intermediate representation can be traversed and analyzed to determine whether they are the second operator.
[0112] The constraint conditions can include that the memory type of the input tensor and / or output tensor of the second operator is the memory on a predetermined execution device. For example, for the gpu::Launch operator, its constraint conditions can be that the memory type of the input tensor and output tensor is GPU memory. If the input tensor and output tensor of this operator are on the CPU side at this time, the tensor data can be copied from the CPU memory to the GPU memory. Among them, the gpu.memcpy_h2d can be called to copy the tensor data from the CPU memory to the GPU memory.
[0113] Step S404: For the read-write dependency relationships between the input tensors and output tensors of each operator in the intermediate representation, enable the input tensors and / or output tensors with read-write dependency relationships to reuse the same memory block on the execution device.
[0114] For example, for operators with a sequential dependency relationship in the topological order of the intermediate representation, such as the output of operator op1 being the input of operator op2, it can be considered that there is a read-write dependency relationship between the output tensor of op1 and the input tensor of op2. Memory reuse processing can be performed on these two tensors, that is, the memory already allocated for the output tensor of op1 and the input tensor of op2 can be rewritten so that they can reuse the same memory area. Thus, the overhead of redundant memory application and data copying can be reduced.
[0115] It should be noted that Figure 4 only one or more of the above steps can be executed, and the execution order of S402 - S404 can also be interchanged. This application does not make any restrictions on this.
[0116] In this way, the memory allocation of the operators in the intermediate representation can be completed, and reference can be returned to Figure 2 :
[0117] Step S203: Generate a target executable file on the execution device based on the result of memory optimization management.
[0118] Among them, the high-level IR in MLIR can be converted into low-level IR through the lower PASS to generate a target executable file on the corresponding hardware platform (i.e., GPU / CPU). The target executable file can be used by the execution device for training or inference based on the neural network model.
[0119] According to the embodiments of this application, by analyzing the intermediate representation corresponding to the computational subgraph on the analysis device and respectively determining the execution device for each operator, corresponding memory optimization management can be performed for different execution devices corresponding to each operator, thereby improving the efficiency of memory application and data copying between execution devices, optimizing the memory space, and reducing the overhead of data transmission. By generating a target executable file on the execution device based on the result of memory optimization management for training or inference based on the neural network model on the execution device, the performance of the model running on the execution device can be improved.
[0120] In the overall process of the memory management method according to the embodiments of the present application, the analysis device can split the running neural network model into subgraphs through the dynamo module under the PyTorch framework to obtain computational subgraphs; the AI compiler on the analysis device can process the operators in the computational subgraphs through the MLIR torch dialect to obtain the multi-level intermediate representation MLIR, and the execution devices of the operators can be determined respectively for each operator in the MLIR (see step S201 for details). For each operator, after determining the execution device, data preprocessing can also be performed. When there is a need to copy discontinuous data between the CPU and the GPU, or between GPUs, the predetermined data parallel processing function in the GPU is used to copy the discontinuous data.
[0121] Then, the low-level IR in the MLIR can also be optimized by means such as constant folding, redundant code elimination, and operator fusion. After the optimization process, according to the execution devices corresponding to the operators determined above, based on the principle of giving priority to running on the GPU, the operators in the optimized IR are repartitioned for the execution devices again to determine the execution devices of each operator respectively.
[0122] After determining the execution devices of each operator, memory optimization management can be performed for different execution devices based on the execution devices corresponding to each operator in the intermediate representation. The process of memory optimization management can be seen in the above Figure 4 shown steps.
[0123] Finally, based on the results of the memory optimization management, a target executable file on the execution device can be generated for the execution device to perform training or inference based on the neural network model (see step S203 above for details).
[0124] Figure 5 shows a schematic diagram of the overall process of the memory management method according to an embodiment of the present application. As Figure 5 shown, the overall process of the memory management method of the present application may include:
[0125] S1, splitting the AI model into subgraphs through the JIT module;
[0126] Among them, the running AI model can be split into subgraphs through the dynamo module under the PyTorch framework to obtain computational subgraphs.
[0127] S2, the AI compiler uses MILR to convert and express the computational subgraph through the torch dialect;
[0128] The operators in the computational subgraph can be transformed and expressed through the MLIR torch dialect to obtain an intermediate representation, namely the multi-level intermediate representation MLIR.
[0129] S3. Perform operator device analysis on the transformed subgraph through the operator device analysis pass.
[0130] The operator device analysis of the transformed subgraph can be performed on a per-operator basis through the MLIR operator device analysis pass module to determine the execution device for executing the operator. This process can be implemented in the manner described in step S201 above.
[0131] S4. Optimize the intermediate representation through passes such as constant folding, redundant code elimination, and operator fusion.
[0132] This process can be implemented in the manner described in step S301 above.
[0133] S5. Apply for memory and synchronize data for the operator device information of the intermediate representation through the operator memory management pass.
[0134] In this operator memory management pass, based on the operator device analysis results in S3, memory allocation and data transfer processing can be performed on the inputs and outputs of the operator to achieve memory application and data synchronization. The data transfer can be implemented in the relevant manner in S201.
[0135] S6. Generate target code on the corresponding hardware platform through the lower pass and execute it.
[0136] The lower pass can generate a target executable file (i.e., target code) on the execution device (i.e., the corresponding hardware platform) based on the results of memory optimization management. Executing the target code can achieve training or inference based on the AI model.
[0137] Figure 6 The structural diagram of a memory management device according to an embodiment of the present application is shown. This device can be used for analysis devices, such as Figure 6 As shown, the device includes:
[0138] The first determination module 601 is configured to determine, for each operator in the intermediate representation, the execution device for executing the operator based on the intermediate representation corresponding to the computational subgraph of the neural network model.
[0139] The optimization management module 602 is configured to perform memory optimization management for different execution devices based on the execution devices corresponding to the operators in the intermediate representation.
[0140] A second determination module 603, configured to generate a target executable file on an execution device based on a result of memory optimization management, where the target executable file is used for the execution device to perform training or inference based on a neural network model.
[0141] In a possible implementation, the first determination module 601 is configured to:
[0142] Determine an execution device for executing an operator according to a memory type of an input tensor and / or an output tensor of the operator indicated in an intermediate representation.
[0143] In a possible implementation, the execution device includes a graphics processing unit (GPU) and a central processing unit (CPU). Determining an execution device for executing an operator according to a memory type of an input tensor and / or an output tensor of the operator indicated in an intermediate representation includes:
[0144] When the memory type of the input tensor and / or the output tensor of the operator is GPU memory, determining that the execution device for executing the operator is the GPU; otherwise,
[0145] When the memory types of the input tensor and the output tensor of the operator are CPU memory, determining that the execution device for executing the operator is the CPU.
[0146] In a possible implementation, the apparatus further includes:
[0147] A copy module, configured to, when there is a need to copy discontinuous data between the CPU and the GPU or between GPUs, use a predetermined data parallel processing function in the GPU to copy the discontinuous data, where the discontinuous data is data that is discontinuously stored in memory units.
[0148] In a possible implementation, the optimization management module 602 is configured to:
[0149] Perform memory pre-allocation under a corresponding execution device based on execution devices corresponding to operators in an intermediate representation.
[0150] In a possible implementation, the optimization management module 602 is configured to:
[0151] For a predetermined first operator in an intermediate representation, cause the first operator to apply for memory in a corresponding execution device outside a scope, where the first operator is an operator within a loop / nested operation, and the scope is the scope of the loop / nested operation.
[0152] In a possible implementation, the optimization management module 602 is configured to:
[0153] For a predetermined second operator in the intermediate representation, based on the constraint conditions corresponding to the second operator, copy the input tensor and / or output tensor of the second operator between execution devices, where the constraint conditions include that the memory type of the input tensor and / or output tensor of the second operator is the memory on the predetermined execution device.
[0154] In a possible implementation, the optimization management module 602 is configured to:
[0155] For the read-write dependency relationships between the input tensors and output tensors among the operators in the intermediate representation, enable the input tensors and / or output tensors with read-write dependency relationships to reuse the same memory block on the execution device.
[0156] In a possible implementation, the device further includes:
[0157] An optimization processing module, configured to perform optimization processing on the intermediate representation to determine the optimized intermediate representation, where the optimization processing includes one or more of constant folding, redundancy elimination, and operator fusion;
[0158] A third determination module, configured to, based on the optimized intermediate representation, respectively determine the execution devices for executing the operators for each operator in the optimized intermediate representation;
[0159] The optimization management module 602 is configured to:
[0160] Based on the execution devices corresponding to the operators in the optimized intermediate representation, perform memory optimization management for different execution devices.
[0161] In a possible implementation, the device can be used in the just-in-time compilation (JIT) stage under the PyTorch framework, and the intermediate representation is a multi-level intermediate representation (MLIR).
[0162] In a possible implementation, the computational subgraph is obtained by splitting the computational graph of the neural network model in the JIT stage, and the intermediate representation is obtained by transforming and expressing the operators of the computational subgraph.
[0163] According to the embodiments of the present application, by respectively determining the execution devices for executing each operator based on the intermediate representation corresponding to the computational subgraph on the analysis device, corresponding memory optimization management can be performed for different execution devices corresponding to each operator, thereby improving the efficiency of memory application and data copying between execution devices, optimizing the memory space, and reducing the overhead of data transmission. By generating the target executable file on the execution device based on the result of the memory optimization management to perform training or inference based on the neural network model on the execution device, the performance of the model running on the execution device can be improved.
[0164] In some embodiments, the functions of the device provided by the embodiments of the present disclosure or the included modules can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0165] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0166] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above methods when executing the instructions stored in the memory.
[0167] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.
[0168] Figure 7 is a block diagram of a device 1900 for memory management shown according to an exemplary embodiment. For example, the device 1900 can be provided as a server. Referring to Figure 7 , the device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above methods.
[0169] The device 1900 may further include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0170] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the device 1900 to complete the above method.
[0171] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0172] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0173] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0174] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0175] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0176] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data - processing apparatus, a device is created that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0177] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0178] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0179] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A memory management method, characterized in that, The method is used for an analysis device and is used in the just-in-time compilation (JIT) phase under the PyTorch framework. The method includes: Based on the intermediate representation corresponding to the computational subgraph of the neural network model, determining the execution device for executing each operator in the intermediate representation respectively. The intermediate representation is a multi-level intermediate representation (MLIR). Based on the execution devices corresponding to the operators in the intermediate representation, performing memory optimization management for different execution devices. Among them, the execution devices are divided based on the principle that the operator preferably runs on the GPU. When the memory type of the input tensor and / or output tensor of the operator in the intermediate representation indicates GPU memory, the execution device corresponding to this operator is the GPU. When the memory type of the input tensor or output tensor of this operator is CPU memory, before performing memory optimization management, the corresponding tensor data is copied from the CPU memory to the GPU memory through a predetermined data parallel processing function in the GPU. In the case where there is a need to copy discontinuous data between the CPU and the GPU, or between GPUs, use the predetermined data parallel processing function in the GPU to copy the discontinuous data. The discontinuous data is data that is not continuously stored in memory units. The performing memory optimization management for different execution devices based on the execution devices corresponding to the operators in the intermediate representation includes: Based on the execution devices corresponding to the operators in the intermediate representation, performing memory pre-allocation for the corresponding execution devices. When the execution device is the GPU, use the gpu.alloc method to allocate GPU memory for the operator. When the execution device is the CPU, use the memerf.alloc method to allocate CPU memory for the operator. For a predetermined second operator in the intermediate representation, based on the second operator satisfying the corresponding constraint conditions, copy the input tensor and / or output tensor of the second operator to the predetermined execution device. The constraint conditions include that the memory type of the input tensor and / or output tensor of the second operator is the memory on the predetermined execution device. Regarding the read-write dependency relationships between the input tensors and output tensors among the operators in the intermediate representation, enable the input tensors and / or output tensors with such read-write dependency relationships to reuse the same piece of memory on the execution device. Among them, when the output tensor of one operator is the input tensor of another operator, it is considered that there is such a read-write dependency relationship between the output tensor of the one operator and the input tensor of the another operator, and rewrite the already allocated memory for the output tensor of the one operator and the input tensor of the another operator to reuse the same piece of memory. Based on the result of the memory optimization management, generate a target executable file on the execution device. The target executable file is used for the execution device to perform training or inference based on the neural network model.
2. The method according to claim 1, wherein The determining the execution device for executing each operator in the intermediate representation based on the intermediate representation corresponding to the computational subgraph of the neural network model includes: Determine the execution device for executing the operator according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation.
3. The method according to claim 2, wherein The execution device includes a graphics processing unit (GPU) and a central processing unit (CPU). The step of determining the execution device for executing the operator according to the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation includes: When the memory type of the input tensor and / or output tensor of the operator is GPU memory, determine that the execution device for executing the operator is the GPU; otherwise, When the memory types of the input tensor and output tensor of the operator are CPU memory, determine that the execution device for executing the operator is the CPU.
4. The method according to any one of claims 1 to 3, characterized in that, The memory optimization management for different execution devices based on the execution devices corresponding to the operators in the intermediate representation includes: For a predetermined first operator in the intermediate representation, cause the first operator to apply for memory in the corresponding execution device outside the scope, where the first operator is an operator within a loop / nested operation, and the scope is the scope of the loop / nested operation.
5. The method according to claim 1, wherein The method further includes: Perform optimization processing on the intermediate representation to determine the optimized intermediate representation. The optimization processing includes one or more of constant folding, redundancy elimination, and operator fusion; Based on the optimized intermediate representation, determine the execution device for executing each operator in the optimized intermediate representation respectively; The memory optimization management for different execution devices based on the execution devices corresponding to the operators in the intermediate representation includes: Based on the execution devices corresponding to the operators in the optimized intermediate representation, perform memory optimization management for different execution devices.
6. The method according to claim 1, characterized in that, The computational subgraph is obtained by splitting the computational graph of the neural network model in the JIT stage, and the intermediate representation is obtained by transforming and expressing the operators of the computational subgraph.
7. A memory management device, characterized in that, The device is used for an analysis device, and the device is used in the just-in-time compilation (JIT) stage under the PyTorch framework. The device includes: A first determination module, configured to determine the execution device for executing each operator in the intermediate representation respectively based on the intermediate representation corresponding to the computational subgraph of the neural network model, where the intermediate representation is a multi-level intermediate representation (MLIR); An optimization management module, configured to perform memory optimization management for different execution devices based on the execution devices corresponding to the operators in the intermediate representation. Among them, the execution devices are divided based on the principle that the operator runs on the GPU first. When the memory type of the input tensor and / or output tensor of the operator indicated in the intermediate representation is GPU memory, the execution device corresponding to the operator is the GPU. When the memory type of the input tensor or output tensor of the operator is CPU memory, before performing memory optimization management, the corresponding tensor data is copied from the CPU memory to the GPU memory through a predetermined data parallel processing function in the GPU. The optimization management module is configured to, when there is a need to copy discontinuous data between the CPU and the GPU or between GPUs, use a predetermined data parallel processing function in the GPU to copy the discontinuous data, where the discontinuous data is data that is discontinuously stored in memory units; The optimization management module is configured to: Based on the execution devices corresponding to the operators in the intermediate representation, perform pre-application of memory under the corresponding execution devices. When the execution device is a GPU, use the gpu.alloc method to apply for GPU memory for the operator. When the execution device is a CPU, use the memerf.alloc method to apply for CPU memory for the operator; For a predetermined second operator in the intermediate representation, based on the second operator satisfying the corresponding constraint conditions, copy the input tensor and / or output tensor of the second operator to a predetermined execution device, where the constraint conditions include that the memory type of the input tensor and / or output tensor of the second operator is the memory on the predetermined execution device; For the read-write dependency relationships between the input tensors and output tensors among the operators in the intermediate representation, enable the input tensors and / or output tensors with the read-write dependency relationships to reuse the same block of memory on the execution device; where, when the output tensor of one operator is the input tensor of another operator, it is considered that there is the read-write dependency relationship between the output tensor of the one operator and the input tensor of the another operator, and rewrite the memory that has been applied for the output tensor of the one operator and the input tensor of the another operator to reuse the same block of memory; A second determination module is configured to generate a target executable file on the execution device based on the result of the memory optimization management, where the target executable file is used for the execution device to perform training or inference based on the neural network model.
8. A memory management device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to, when executing the instructions stored in the memory, implement the method according to any one of claims 1 to 6.
9. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by the processor, implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Memory management method, system and equipment and computer readable storage medium
CN114816752A