A GPU memory optimization management method and system
By dividing the memory space on the GPU device side into multiple areas and formulating a memory allocation plan, the problem of frequent alternating operation of GPU and CPU programs in the heterogeneous programming framework is solved, and the efficiency of neural network model inference is improved.
Patent Information
- Application Number
- CN202510188597.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-20
AI Technical Summary
During the neural network model inference process, the existing heterogeneous programming framework causes programs to frequently alternate between the GPU device side of the graphics processor and the CPU host side of the CPU, reducing the efficiency of model inference.
By dividing the memory space on the GPU device side of the graphics processor into instruction memory area, controlled memory area and autonomous memory area, and formulating a memory allocation plan, the utilization of memory space and memory reuse are optimized.
The utilization of memory space is optimized, the memory reuse is maximized, the number of alternate executions of CPU and GPU programs is reduced, and the efficiency of neural network model inference is improved.
Smart Images

Figure CN119645672B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of memory management, and in particular to a GPU memory optimization management method and system. Background Art
[0002] With the development of deep learning technology, neural network models are being used more and more. In order to achieve better learning results, neural network models are being designed to be more complex and the model depth is getting deeper. The process of model reasoning involves a large number of data calculation tasks. Therefore, in practical applications, heterogeneous programming is often required. After the host CPU obtains the input data, the parallel computing power of the graphics processor GPU device is used to perform data calculation tasks to accelerate the process of neural network model reasoning.
[0003] Most of the existing mainstream heterogeneous programming frameworks manage the GPU device-side memory on the CPU host side, and dynamically allocate memory space for each layer of the neural network in the device-side memory area in sequence according to the order in which the network memory is used. When the data stored in the space is no longer used by subsequent programs, the memory is released in time to avoid memory overflow. This method easily causes the GPU device-side program and the CPU host-side program to run alternately many times during the entire neural model reasoning process. The more complex the neural network is, the more times it needs to run alternately, which greatly reduces the efficiency of model reasoning.
[0004] Based on the above problems, the present invention proposes a GPU memory optimization management method and system. Summary of the invention
[0005] In order to overcome the defects of the prior art, the present invention provides a simple and efficient GPU memory optimization management method and system.
[0006] The present invention is achieved through the following technical solutions:
[0007] A GPU memory optimization management method, characterized in that: it is applicable to a heterogeneous computing platform, the hardware architecture of the heterogeneous computing platform is divided into a host side and a device side, and the two are interconnected through a PCIe high-speed bus;
[0008] The following steps are involved:
[0009] Step S1, dividing the device into regions;
[0010] Based on the data scale processed by the neural network and the program scale of the neural network itself, the memory space of the graphics processor GPU device is divided into three areas: instruction memory area, controlled memory area and autonomous memory area;
[0011] The instruction memory area is used to store the executable program generated by compiling the kernel functions of each layer of the network during neural network inference, as well as the model weights and bias parameters required for model inference;
[0012] A controlled memory area is used to store the raw input data that the neural network needs to process, the processed output data, and the data structure that stores the memory allocation scheme used during the model inference process;
[0013] Autonomous memory area, which is used autonomously by the GPU device program according to the memory allocation plan, and is used to allocate memory space for the input and output data required by the kernel functions of each layer of the network during the inference process;
[0014] Step S2: Formulate a memory allocation plan during the neural network model reasoning process;
[0015] According to the network hierarchy of the neural network model, the corresponding GPU device-side kernel function is customized and designed for each layer of the network;
[0016] Analyze the model reasoning process, determine the execution order of each kernel function, and construct a directed acyclic graph with kernel functions as nodes;
[0017] According to the input and output parameter data of each kernel function and the data dependency relationship between kernel functions in the directed acyclic graph, the life cycle of each memory object is determined, so as to further determine the memory space occupied by each kernel function when it is executed, and generate a kernel function memory requirement table;
[0018] Determine the size of the memory space to be applied for based on the kernel function memory requirement table. According to the memory requirements of each kernel function in the kernel function memory requirement table, consider the memory space whose life cycle has ended during the allocation process to generate the starting address and size of each memory space.
[0019] In step S2, the kernel function with the largest memory space requirement is searched in the kernel function memory requirement table, and the memory space size of the kernel function is used as the memory space size applied for by the entire neural network model in the autonomous memory area.
[0020] Step S3, writing the memory allocation plan;
[0021] Construct a memory management container, assign a key value to each kernel function according to the hierarchical structure of the specific neural network model, and assign a single linked list to each key value, the single linked list is used to store the starting address and size of all memory spaces required by the corresponding kernel function generated in step S2;
[0022] After completing the data filling of the memory management container, the host program writes the memory manager data structure into the specified memory of the device-side controlled memory area;
[0023] In step S3, the parameters of the memory management container include the source data address, the destination data address, the base address of the model reasoning process data and the kernel function memory allocation subcontainer (key-value pair).
[0024] Step S4, executing the memory allocation scheme;
[0025] When the neural network model is executed on the device side, it first reads the parameters of the memory management container from the specified memory in the controlled memory area, and queries the memory manager to obtain the memory space required for each kernel function in the reasoning process through the predetermined key value and the single linked list; until the entire reasoning process is completed, the reasoning result data is stored in the output data memory space specified in the controlled memory area.
[0026] A GPU memory optimization management system, comprising a region division module, a memory allocation plan formulation module, a memory allocation plan writing module and a memory allocation plan execution module;
[0027] The area division module is responsible for dividing the memory space of the graphics processor GPU device into three areas based on the data scale processed by the neural network and the program scale of the neural network itself: the instruction memory area, the controlled memory area and the autonomous memory area;
[0028] The instruction memory area is used to store the executable program generated by compiling the kernel functions of each layer of the network during neural network inference, as well as the model weights and bias parameters required for model inference;
[0029] A controlled memory area is used to store the raw input data that the neural network needs to process, the processed output data, and the data structure that stores the memory allocation scheme used during the model inference process;
[0030] Autonomous memory area, which is used autonomously by the GPU device program according to the memory allocation plan, and is used to allocate memory space for the input and output data required by the kernel functions of each layer of the network during the inference process;
[0031] The memory allocation plan formulation module is responsible for formulating the memory allocation plan during the neural network model inference process; the specific process is as follows:
[0032] Step S2.1, according to the network hierarchy of the neural network model, custom design the corresponding GPU device-side kernel function for each layer of the network;
[0033] Step S2.2, analyzing the model reasoning process, determining the execution order of each kernel function, and constructing a directed acyclic graph with the kernel function as a node;
[0034] Step S2.3, according to the input and output parameter data of each kernel function and the data dependency relationship between kernel functions in the directed acyclic graph, determine the life cycle of each memory object, thereby further determining the memory space occupied by each kernel function when executing, and generating a kernel function memory requirement table;
[0035] Step S2.4, determine the size of the memory space to be applied for based on the kernel function memory requirement table, and according to the memory requirements of each kernel function in the kernel function memory requirement table, consider reusing the memory space whose life cycle has ended during the allocation process, and generate the starting address and size of each memory space;
[0036] The memory allocation scheme formulation module searches for the kernel function with the largest memory space requirement in the kernel function memory requirement table, and uses the memory space size of the kernel function as the memory space size applied for by the entire neural network model in the autonomous memory area.
[0037] The memory allocation scheme writing module is responsible for constructing a memory management container. According to the hierarchical structure of the specific neural network model, a key value is assigned to each kernel function, and a single linked list is assigned to each key value. The single linked list is used to store the starting address and size of all memory spaces required by the corresponding kernel function generated in the memory allocation scheme formulation module.
[0038] After completing the data filling of the memory management container, the memory manager data structure is written into the specified memory of the controlled memory area on the device side through the host side program;
[0039] The parameters of the memory management container include a source data address, a destination data address, a base address of the model reasoning process data, and a kernel function memory allocation subcontainer (key-value pair).
[0040] The memory allocation scheme execution module is responsible for executing the neural network model on the device side. It first reads the parameters of the memory management container from the specified memory in the controlled memory area, and queries the memory manager to obtain the memory space required for each kernel function in the reasoning process through the predetermined key value and single linked list until the entire reasoning process is completed, and stores the reasoning result data into the output data memory space specified in the controlled memory area.
[0041] A GPU memory optimization management device, characterized by comprising:
[0042] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.
[0043] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and the computer program implements the above method when executed by a processor.
[0044] The beneficial effects of the present invention are: the GPU memory optimization management method and system optimize the utilization of memory space, achieve maximum memory reuse, avoid frequent alternating execution of programs on the central processing unit (CPU) host side and the graphics processing unit (GPU) device side, reduce the overall scheduling overhead of the system, and improve reasoning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Attached Figure 1 Schematic diagram of the GPU memory optimization management method of the present invention.
[0047] Attached Figure 2 This is a schematic diagram of DDR memory partitioning on the device side of the present invention.
[0048] Attached Figure 3 Schematic diagram of the neural network model hierarchy of the present invention.
[0049] Attached Figure 4 Schematic diagram of the execution sequence of the kernel function of the present invention. DETAILED DESCRIPTION
[0050] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0051] The GPU memory optimization management method is applicable to heterogeneous computing platforms. The hardware architecture of the heterogeneous computing platform is divided into two parts: the host side and the device side, which are connected to each other through a PCIe high-speed bus.
[0052] The following steps are involved:
[0053] Step S1, dividing the device into regions;
[0054] Based on the data scale processed by the neural network and the program scale of the neural network itself, the memory space of the graphics processor GPU device is divided into three areas: instruction memory area, controlled memory area and autonomous memory area;
[0055] The instruction memory area is used to store the executable program generated by compiling the kernel functions of each layer of the network during neural network inference, as well as the model weights and bias parameters required for model inference;
[0056] A controlled memory area is used to store the raw input data that the neural network needs to process, the processed output data, and the data structure that stores the memory allocation scheme used during the model inference process;
[0057] Autonomous memory area, which is used autonomously by the GPU device program according to the memory allocation plan, and is used to allocate memory space for the input and output data required by the kernel functions of each layer of the network during the inference process;
[0058] Step S2: Formulate a memory allocation plan during the neural network model reasoning process;
[0059] As attached Figure 3 As shown, according to the network hierarchy of the neural network model, the corresponding GPU device-side kernel function is custom designed for each layer of the network;
[0060] As attached Figure 4 As shown, the model reasoning process is analyzed, the execution order of each kernel function is determined, and a directed acyclic graph is constructed with the kernel function as the node;
[0061] According to the input and output parameter data of each kernel function and the data dependency relationship between kernel functions in the directed acyclic graph, the life cycle of each memory object is determined, so as to further determine the memory space occupied by each kernel function when it is executed, and generate the kernel function memory requirement table, as shown in the following table;
[0062] Table 1 Kernel function memory requirements
[0063]
[0064] Determine the size of the memory space to be applied for based on the kernel function memory requirement table. According to the memory requirements of each kernel function in the kernel function memory requirement table, consider the memory space whose life cycle has ended during the allocation process to generate the starting address and size of each memory space.
[0065] In step S2, the kernel function with the largest memory space requirement is searched in the kernel function memory requirement table, and the memory space size of the kernel function is used as the memory space size applied for by the entire neural network model in the autonomous memory area.
[0066] Find the kernel function with the largest memory space requirement in Table 1, fun6, which occupies 25M memory space. Then use this memory space size (25M) as the memory space size applied for by the entire neural network model in the autonomous memory area.
[0067] To further explain the 25M memory allocation during neural network inference, the following detailed timeline and table can be used to deduce the memory allocation:
[0068] Table 2 Memory allocation scheme table
[0069]
[0070] Step S3, writing the memory allocation plan;
[0071] Construct a memory management container, assign a key value to each kernel function according to the hierarchical structure of the specific neural network model, and assign a single linked list to each key value, the single linked list is used to store the starting address and size of all memory spaces required by the corresponding kernel function generated in step S2;
[0072] After completing the data filling of the memory management container, the host program writes the memory manager data structure into the specified memory of the device-side controlled memory area;
[0073] In step S3, the parameters of the memory management container include the source data address, the destination data address, the base address of the model reasoning process data and the kernel function memory allocation subcontainer (key-value pair).
[0074] Step S4, executing the memory allocation scheme;
[0075] When the neural network model is executed on the device side, it first reads the parameters of the memory management container from the specified memory in the controlled memory area, and queries the memory manager to obtain the memory space required for each kernel function in the reasoning process through the predetermined key value and the single linked list; until the entire reasoning process is completed, the reasoning result data is stored in the output data memory space specified in the controlled memory area.
[0076] The GPU memory optimization management system includes a region division module, a memory allocation plan formulation module, a memory allocation plan writing module and a memory allocation plan execution module;
[0077] The area division module is responsible for dividing the memory space of the graphics processor GPU device into three areas based on the data scale processed by the neural network and the program scale of the neural network itself: the instruction memory area, the controlled memory area and the autonomous memory area;
[0078] The instruction memory area is used to store the executable program generated by compiling the kernel functions of each layer of the network during neural network inference, as well as the model weights and bias parameters required for model inference;
[0079] A controlled memory area is used to store the raw input data that the neural network needs to process, the processed output data, and the data structure that stores the memory allocation scheme used during the model inference process;
[0080] Autonomous memory area, which is used autonomously by the GPU device program according to the memory allocation plan, and is used to allocate memory space for the input and output data required by the kernel functions of each layer of the network during the inference process;
[0081] The memory allocation plan formulation module is responsible for formulating the memory allocation plan during the neural network model inference process; the specific process is as follows:
[0082] Step S2.1, according to the network hierarchy of the neural network model, custom design the corresponding GPU device-side kernel function for each layer of the network;
[0083] Step S2.2, analyzing the model reasoning process, determining the execution order of each kernel function, and constructing a directed acyclic graph with the kernel function as a node;
[0084] Step S2.3, according to the input and output parameter data of each kernel function and the data dependency relationship between kernel functions in the directed acyclic graph, determine the life cycle of each memory object, thereby further determining the memory space occupied by each kernel function when executing, and generating a kernel function memory requirement table;
[0085] Step S2.4, determine the size of the memory space to be applied for based on the kernel function memory requirement table, and according to the memory requirements of each kernel function in the kernel function memory requirement table, consider reusing the memory space whose life cycle has ended during the allocation process, and generate the starting address and size of each memory space;
[0086] The memory allocation scheme formulation module searches for the kernel function with the largest memory space requirement in the kernel function memory requirement table, and uses the memory space size of the kernel function as the memory space size applied for by the entire neural network model in the autonomous memory area.
[0087] The memory allocation scheme writing module is responsible for constructing a memory management container, assigning a key value to each kernel function according to the hierarchical structure of the specific neural network model, and assigning a single linked list to each key value. The single linked list is used to store the starting address and size of all memory spaces required by the corresponding kernel function generated in step S2;
[0088] After completing the data filling of the memory management container, the memory manager data structure is written into the specified memory of the controlled memory area on the device side through the host side program;
[0089] The parameters of the memory management container include a source data address, a destination data address, a base address of the model reasoning process data, and a kernel function memory allocation subcontainer (key-value pair).
[0090] The memory allocation scheme execution module is responsible for executing the neural network model on the device side. It first reads the parameters of the memory management container from the specified memory in the controlled memory area, and queries the memory manager to obtain the memory space required for each kernel function in the reasoning process through the predetermined key value and single linked list until the entire reasoning process is completed, and stores the reasoning result data into the output data memory space specified in the controlled memory area.
[0091] The GPU memory optimization management device includes:
[0092] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.
[0093] The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0094] In summary, the GPU memory optimization management method and system analyzes the memory usage of each kernel function according to the hierarchical structure of the neural network, calculates the peak space of memory usage during the inference process, and pre-allocates the memory usage plan of each kernel function according to the peak space;
[0095] The memory allocation plan is stored in the form of a memory management container data structure and written to the controlled memory area; the memory allocation plan is obtained from the controlled memory area once during device execution, and the location of the memory required by each kernel function is obtained by querying the memory allocation plan information during the execution process; the frequent alternating execution of the central processing unit CPU host side and the graphics processing unit GPU device side program is avoided, thereby improving the inference efficiency;
[0096] In the process of formulating the memory allocation plan, the utilization of memory space is optimized, the memory life cycle is taken into consideration, and maximum memory reuse is achieved.
[0097] By specifying the memory allocation plan on the host side and executing it on the device side, frequent memory application and release in the system is avoided, thus reducing the overall scheduling overhead of the system.
[0098] The embodiment described above is only one specific implementation of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A GPU memory optimization management method, characterized by: Applicable to heterogeneous computing platforms, the hardware architecture of which is divided into two parts: the host side and the device side, which are interconnected through the PCIe high-speed bus; including the following steps: Step S1, dividing the device into regions; Based on the data scale processed by the neural network and the program scale of the neural network itself, the memory space of the graphics processor GPU device is divided into three areas: instruction memory area, controlled memory area and autonomous memory area; Step S2: Formulate a memory allocation plan during the neural network model reasoning process; According to the network hierarchy of the neural network model, the corresponding GPU device-side kernel function is customized and designed for each layer of the network; Analyze the model reasoning process, determine the execution order of each kernel function, and construct a directed acyclic graph with kernel functions as nodes; According to the input and output parameter data of each kernel function and the data dependency relationship between kernel functions in the directed acyclic graph, the life cycle of each memory object is determined, so as to further determine the memory space occupied by each kernel function when it is executed, and generate a kernel function memory requirement table; Determine the size of the memory space to be applied for based on the kernel function memory requirement table. According to the memory requirements of each kernel function in the kernel function memory requirement table, consider the memory space whose life cycle has ended during the allocation process to generate the starting address and size of each memory space. Step S3, writing the memory allocation plan; Construct a memory management container, assign a key value to each kernel function according to the hierarchical structure of the specific neural network model, and assign a single linked list to each key value, the single linked list is used to store the starting address and size of all memory spaces required by the corresponding kernel function generated in step S2; After completing the data filling of the memory management container, the host program writes the memory manager data structure into the specified memory of the device-side controlled memory area; Step S4, executing the memory allocation scheme; When the neural network model is executed on the device side, it first reads the parameters of the memory management container from the specified memory in the controlled memory area, and queries the memory manager to obtain the memory space required for each kernel function in the reasoning process through the predetermined key value and the single linked list; until the entire reasoning process is completed, the reasoning result data is stored in the output data memory space specified in the controlled memory area.
2. The GPU memory optimization management method according to claim 1, characterized in that: The instruction memory area is used to store the executable program generated after the kernel function of each layer of the network is compiled during neural network reasoning, as well as the model weights and bias parameters required for model reasoning; A controlled memory area is used to store the raw input data that the neural network needs to process, the processed output data, and the data structure that stores the memory allocation scheme used during the model inference process; The autonomous memory area is used autonomously by the graphics processor GPU device program according to the memory allocation plan. It is used to allocate memory space for the input and output data required by the kernel functions of each layer of the network during the inference process.
3. The GPU memory optimization management method according to claim 1, characterized in that: In step S2, the kernel function with the largest memory space requirement is searched in the kernel function memory requirement table, and the memory space size of the kernel function is used as the memory space size applied for by the entire neural network model in the autonomous memory area.
4. The GPU memory optimization management method according to claim 1, characterized in that: In step S3, the parameters of the memory management container include the source data address, the destination data address, the base address of the model reasoning process data and the kernel function memory allocation subcontainer.
5. A GPU memory optimization management system, characterized by: It includes a region division module, a memory allocation plan formulation module, a memory allocation plan writing module and a memory allocation plan execution module; The area division module is responsible for dividing the memory space of the graphics processor GPU device into three areas based on the data scale processed by the neural network and the program scale of the neural network itself: the instruction memory area, the controlled memory area and the autonomous memory area; The memory allocation plan formulation module is responsible for formulating the memory allocation plan during the neural network model inference process; the specific process is as follows: Step S2.1, according to the network hierarchy of the neural network model, custom design the corresponding GPU device-side kernel function for each layer of the network; Step S2.2, analyzing the model reasoning process, determining the execution order of each kernel function, and constructing a directed acyclic graph with the kernel function as a node; Step S2.3, according to the input and output parameter data of each kernel function and the data dependency relationship between kernel functions in the directed acyclic graph, determine the life cycle of each memory object, thereby further determining the memory space occupied by each kernel function when executing, and generating a kernel function memory requirement table; Step S2.4, determine the size of the memory space to be applied for based on the kernel function memory requirement table, and according to the memory requirements of each kernel function in the kernel function memory requirement table, consider reusing the memory space whose life cycle has ended during the allocation process, and generate the starting address and size of each memory space; The memory allocation scheme writing module is responsible for constructing a memory management container. According to the hierarchical structure of the specific neural network model, a key value is assigned to each kernel function, and a single linked list is assigned to each key value. The single linked list is used to store the starting address and size of all memory spaces required by the corresponding kernel function generated in the memory allocation scheme formulation module. After completing the data filling of the memory management container, the memory manager data structure is written into the specified memory of the controlled memory area on the device side through the host side program; The memory allocation scheme execution module is responsible for executing the neural network model on the device side. It first reads the parameters of the memory management container from the specified memory in the controlled memory area, and queries the memory manager to obtain the memory space required for each kernel function in the reasoning process through the predetermined key value and single linked list until the entire reasoning process is completed, and stores the reasoning result data into the output data memory space specified in the controlled memory area.
6. The GPU memory optimization management system according to claim 5, characterized in that: The instruction memory area is used to store the executable program generated after the kernel function of each layer of the network is compiled during neural network reasoning, as well as the model weights and bias parameters required for model reasoning; A controlled memory area is used to store the raw input data that the neural network needs to process, the processed output data, and the data structure that stores the memory allocation scheme used during the model inference process; The autonomous memory area is used autonomously by the graphics processor GPU device program according to the memory allocation plan. It is used to allocate memory space for the input and output data required by the kernel functions of each layer of the network during the inference process.
7. The GPU memory optimization management system according to claim 5, characterized in that: The memory allocation scheme formulation module searches for the kernel function with the largest memory space requirement in the kernel function memory requirement table, and uses the memory space size of the kernel function as the memory space size applied for by the entire neural network model in the autonomous memory area.
8. The GPU memory optimization management system according to claim 5, characterized in that: The parameters of the memory management container include a source data address, a destination data address, a base address of the model reasoning process data, and a kernel function memory allocation subcontainer.
9. A GPU memory optimization management device, characterized in that: include: One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 4.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Data processing method, model optimization device and model execution device
CN112529169A
Memory management method and device and storage medium
CN117667424A