Memory allocation method for deep learning model, computer equipment and medium
By reasonably allocating graphics card memory based on the model parameter quantity, input size and expected maximum batch volume, the problem of unreasonable allocation of large deep learning models on graphics card memory is solved, and more efficient memory utilization and model inference speed are achieved.
Patent Information
- Application Number
- CN202510346473.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The memory allocation of large-scale deep learning models in the prior art on graphics cards is unreasonable, which can easily lead to memory overflow or a decrease in model inference speed.
By obtaining the model parameter quantity, input size and expected maximum batch quantity of the target deep learning model, the expected memory capacity is determined, and the graphics card memory is reasonably allocated according to this capacity, including dynamic storage of parameter sets, memory sharing of the same weight matrix, decomposition of matrix multiplication operations, and adjusting the maximum batch quantity, the allocation of graphics card memory resources is optimized.
It reduces memory overflow and model inference speed, improves the efficiency of graphics card memory utilization and model inference speed.
Smart Images

Figure CN120276844A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of memory allocation, and specifically relates to a memory allocation method, a computer device, and a medium for a deep learning model. Background Art
[0002] With the rapid development of deep learning and natural language processing technologies, large pre-trained models (large models) occupy an important position in the field of artificial intelligence due to their excellent generalization ability and multi-task processing ability. However, with the continuous expansion of model scale, significant technical challenges have been encountered in their actual deployment and application. These challenges not only limit the application scope of large models but also pose higher requirements for building an efficient and scalable large model development platform.
[0003] When configuring the video memory for large models, relevant technical personnel often manually allocate the corresponding video card memory size based on past experience. However, the video card memory allocated in this way is often not reasonable enough, and when the video card memory consumed by large models is excessive, problems such as memory overflow or a decrease in model inference speed are likely to occur. Summary of the Invention
[0004] Embodiments of this application provide a memory allocation method, a computer device, and a medium for a deep learning model, aiming to make the video card memory allocation of the target deep learning model more reasonable to reduce the situation of memory overflow or a decrease in model inference speed.
[0005] In a first aspect, embodiments of this application provide a memory allocation method for a deep learning model. The memory allocation method for the deep learning model includes:
[0006] Obtain the number of model parameters and the input size in the target deep learning model;
[0007] Obtain the expected maximum batch size of the target deep learning model;
[0008] Based on the number of model parameters, the input size, and the expected maximum batch size, determine the expected memory capacity of the target deep learning model;
[0009] Allocate video card memory for the target deep learning model according to the expected memory capacity.
[0010] In some embodiments, after obtaining the number of model parameters in the target deep learning model, it further includes:
[0011] Divide the model parameters in the target deep learning model into multiple parameter sets;
[0012] After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes:
[0013] During the model inference stage of the target deep learning model, determine the requested model parameters;
[0014] Store the parameter set where the requested model parameters are located in the video card memory of the target deep learning model.
[0015] In some embodiments, after storing the parameter set where the requested model parameters are located in the video card memory of the target deep learning model, it further includes:
[0016] If the requested model parameters have been used up, determine the parameter type of the used-up model parameters;
[0017] If the parameter type is a preset parameter type, transfer the used-up model parameters to the memory of the host where the video card is located for storage.
[0018] In some embodiments, the target deep learning model includes a multi-layer network structure connected in sequence. After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes:
[0019] In the multi-layer network structure connected in sequence, determine at least two network structures with the same weight matrix;
[0020] In the video card memory of the target deep learning model, simultaneously allocate the memory area storing the same weight matrix to the at least two network structures.
[0021] In some embodiments, the target deep learning model includes a multi-layer network structure connected in sequence. After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes:
[0022] During the model inference stage of the target deep learning model, obtain the model inference result of the target layer network structure where the model inference has ended, so as to use the model inference result for the model inference of the next layer network structure of the target layer network structure;
[0023] In the video card memory of the target deep learning model, delete the intermediate data of the model inference process of the target layer network structure, and the intermediate data is used to determine the model inference result.
[0024] In some embodiments, at least one layer of network structure in the target deep learning model includes matrix multiplication operations. After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes:
[0025] During the model inference phase of the target deep learning model, decompose the matrix multiplication operation into multiple sub-operations;
[0026] Utilize the video card memory of the target deep learning model to sequentially execute multiple sub-operations.
[0027] In some embodiments, the method further includes:
[0028] After each sub-operation is executed, perform sequential accumulation processing on the result of the executed sub-operation to obtain the result of the matrix multiplication operation.
[0029] In some embodiments, after determining the expected memory capacity of the target deep learning model based on the number of model parameters, the input size, and the expected maximum batch size, it further includes:
[0030] If the difference between the remaining capacity of the video card memory and the expected memory capacity is less than a preset difference, reduce the expected maximum batch size according to the remaining capacity of the video card memory to adjust the expected memory capacity;
[0031] After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes:
[0032] Perform model inference on the target deep learning model according to the reduced expected maximum batch size.
[0033] In a second aspect, an embodiment of the present application provides a memory allocation device for a deep learning model, and the memory allocation device for the deep learning model includes:
[0034] A first acquisition module for acquiring the number of model parameters and the input size in the target deep learning model;
[0035] A second acquisition module for acquiring the expected maximum batch size of the target deep learning model;
[0036] A memory determination module for determining the expected memory capacity of the target deep learning model based on the number of model parameters, the input size, and the expected maximum batch size;
[0037] A memory allocation module for allocating video card memory for the target deep learning model according to the expected memory capacity.
[0038] In a third aspect, an embodiment of the present application provides a computer device, which includes a processor and a memory, and a computer program is stored in the memory, and the computer program is configured to be executed by the processor to implement the memory allocation method for a deep learning model as described in any one of the above.
[0039] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program configured to be executed by a processor to implement the memory allocation method for a deep learning model as described in any one of the above.
[0040] Fifthly, an embodiment of the present application provides a computer program product including a computer program or instruction, where the computer program or instruction is executed by a processor to implement the memory allocation method for a deep learning model as described in any one of the above.
[0041] Beneficial effects of the embodiments of the present application:
[0042] In the embodiments of the present application, by determining the expected memory capacity of the target deep learning model based on the number of model parameters, input size, and expected memory capacity in the target deep learning model, the video card memory allocation of the target deep learning model is made more reasonable to reduce the situation of memory overflow or the decrease in model inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 is a schematic flowchart of an embodiment of the memory allocation method for a deep learning model provided by the embodiment of the present application;
[0045] Figure 2 is a schematic flowchart of another embodiment of the memory allocation method for a deep learning model provided by the embodiment of the present application;
[0046] Figure 3 is a schematic flowchart of still another embodiment of the memory allocation method for a deep learning model provided by the embodiment of the present application;
[0047] Figure 4 is a schematic flowchart of yet another embodiment of the memory allocation method for a deep learning model provided by the embodiment of the present application;
[0048] Figure 5 is a schematic structural diagram of an embodiment of a computer device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0050] In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined. In addition, in the description of the present application, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features.
[0051] In order to make the video card memory allocation of the target deep learning model more reasonable, the embodiments of the present application provide a memory allocation method, a computer device and a medium for a deep learning model. By determining the expected memory capacity of the target deep learning model based on the number of model parameters, input size and expected memory capacity in the target deep learning model, the video card memory allocation of the target deep learning model is made more reasonable to reduce the situation of memory overflow or the decrease in model inference speed. For the specific solution, please refer to the following specific description.
[0052] In the first aspect, the embodiments of the present application provide a memory allocation method for a deep learning model. Specifically, refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the memory allocation method for a deep learning model. In Figure 1 , the memory allocation method for a deep learning model may include:
[0053] 101. Obtain the number of model parameters and the input size in the target deep learning model.
[0054] In an embodiment of the present application, the target deep learning model may be a deep learning model that requires post-training reinforcement inference. Post-training is relative to pre-training. By performing post-training on the target deep learning model, the inference ability of the target deep learning model can be strengthened. The target deep learning model that requires post-training reinforcement inference may be, for example, a large pre-trained model (large model). The model parameter quantity refers to the number of model parameters. The input size refers to the size of the input data of the target deep learning model. For example, it may include: the batch size and input feature number of the fully connected layer in the target deep learning model, the batch size, input channel number, height, and width of the convolutional layer, the batch size, input channel number, height, and width of the pooling layer, the batch size and input feature number of the normalization layer, and the batch size and input feature number of the activation layer.
[0055] 102. Obtain the expected maximum batch processing volume of the target deep learning model.
[0056] In an embodiment of the present application, the expected maximum batch processing volume refers to the expected maximum batch size of the target deep learning model. The expected maximum batch processing volume can be set within a preset numerical range based on actual requirements.
[0057] 103. Based on the model parameter quantity, input size, and expected maximum batch processing volume, determine the expected memory capacity of the target deep learning model.
[0058] In an embodiment of the present application, since when the model parameter quantity, input size, and expected maximum batch processing volume are larger, the target deep learning model itself is more complex, and the graphics card memory required for model calculation is usually more. Therefore, based on previous experiments, a better correlation relationship between the model parameter quantity, input size, expected maximum batch processing volume, and expected memory capacity can be generated. Then, based on this correlation relationship, the expected memory capacity of the target deep learning model can be determined.
[0059] 104. Allocate graphics card memory for the target deep learning model according to the expected memory capacity.
[0060] In an embodiment of the present application, during the memory initialization stage of the target deep learning model, first allocate graphics card memory for the target deep learning model according to the expected memory capacity. The capacity of the graphics card memory allocated to the target deep learning model can be equal to the expected memory capacity, so that the allocation of the graphics card memory of the target deep learning model is more reasonable.
[0061] It can be seen that in the above embodiments of the present application, by determining the expected memory capacity of the target deep learning model based on the number of model parameters, input size, and expected memory capacity in the target deep learning model, that is, the above embodiments of the present application provide a memory optimization solution in the initialization stage to dynamically allocate the video card memory resources: estimate the required video memory (i.e., the expected memory capacity) according to the model parameters, input size, and expected maximum batch size, so that the video card memory allocation of the target deep learning model is more reasonable to reduce the situation of memory overflow or the decline of model inference speed. In addition, in the step of determining the expected memory capacity of the target deep learning model, the number of concurrent tasks of the target deep learning model can also be comprehensively considered to ensure that each inference instance can obtain sufficient resources.
[0062] In addition, to prevent memory shortage in case of emergencies, about 10% of the video memory can be reserved as a buffer in the video card. This can be used to handle unexpectedly increased temporary variables or more complex dynamic computational graph generation requirements.
[0063] In some embodiments of the present application, as Figure 2 shown, on the basis of the embodiment shown in Figure 1 in order to reduce the video card memory occupied by the target deep learning model, after allocating the video card memory for the target deep learning model according to the expected memory capacity, it may further include:
[0064] 201. In the model inference stage of the target deep learning model, determine the requested model parameters.
[0065] In the embodiments of the present application, when initializing the memory of the target deep learning model, the target deep learning model can enter the model inference stage. In the model inference stage of the target deep learning model, the model parameters of the target deep learning model need to be requested for model calculation in the video card, and at this time, the requested model parameters can be determined.
[0066] 202. Store the parameter set where the requested model parameters are located into the video card memory of the target deep learning model.
[0067] In an embodiment of the present application, the model parameters in the target deep learning model are divided into multiple parameter sets, and each parameter set includes a part of the model parameters in the target deep learning model. Since the model parameters in the target deep learning model are stored in the memory of the host where the graphics card is located, after determining the requested model parameters, the parameter set where the requested model parameters are located can be determined in the memory of the host where the graphics card is located, and then the parameter set where the requested model parameters are located is stored in the graphics card memory of the target deep learning model for model calculation in the graphics card. The parameter sets where the unrequested model parameters are located remain stored in the memory of the host where the graphics card is located and are not stored in the graphics card memory of the target deep learning model to reduce the occupancy of the graphics card memory by the target deep learning model.
[0068] It can be seen that in the above embodiment of the present application, the model parameters in the target deep learning model are divided into multiple parameter sets, and only when the model parameters of a specific parameter set are requested, they are migrated from the memory of the host memory to the video memory, thus significantly reducing the video memory occupancy.
[0069] In some embodiments of the present application, an example of the division strategy for multiple parameter sets is described. Specifically, dividing the model parameters in the target deep learning model into multiple parameter sets may include: obtaining the remaining capacity of the graphics card memory after allocating the graphics card memory for the target deep learning model according to the expected memory capacity; determining the target size of the parameter set corresponding to the remaining capacity of the graphics card memory, where the corresponding relationship between the remaining capacity of the graphics card memory and the target size of the parameter set can be obtained based on previous experiments and is not limited herein; dividing the model parameters in the target deep learning model into multiple parameter sets according to the target size of the parameter set, so that the size of each parameter set is less than the target size of the parameter set to avoid insufficient graphics card memory in case of emergencies and further reduce the occurrence of memory overflow.
[0070] In some embodiments of the present application, after step 202, it may further include: if the requested model parameters are no longer in use, determining the parameter type of the model parameters that are no longer in use; if the parameter type is a preset parameter type, transferring the model parameters that are no longer in use to the memory of the host where the graphics card is located for storage, so as to avoid the model parameters that are no longer in use from continuing to occupy the graphics card memory, thereby further reducing the video memory occupancy. The preset parameter type can be set in advance (that is, for some models, it can be predefined which model parameters can be safely moved back to the memory of the host when they are no longer in use, thus saving the video memory).
[0071] In some embodiments of the present application, since the target deep learning model includes a multi-layer network structure connected in sequence, different layers of the network structure can also be simplified to further reduce the video memory occupation. Specifically, after allocating video card memory for the target deep learning model according to the expected memory capacity, it may further include: in the multi-layer network structure connected in sequence, determining at least two layers of network structures with the same weight matrix; in the video card memory of the target deep learning model, simultaneously allocating the memory area storing the same weight matrix to at least two layers of network structures, that is, at least two layers of network structures with the same weight matrix can share the same memory area for model calculation without separately allocating a memory area, so as to reduce the video memory occupation (for example, if two adjacent layers of network structures have weight matrices of the same size, the memory area already allocated for the previous layer can be directly used without reallocating a new memory area).
[0072] In some embodiments of the present application, in order to further reduce the video memory occupation, after allocating video card memory for the target deep learning model according to the expected memory capacity, it may further include: in the model inference stage of the target deep learning model, obtaining the model inference result of the target layer network structure where the model inference has ended, so as to use the model inference result for the model inference of the next layer network structure of the target layer network structure; in the video card memory of the target deep learning model, deleting the intermediate data of the model inference process of the target layer network structure, and the intermediate data is used to determine the model inference result. It can be understood that when the model inference of the target layer network structure has ended, the intermediate data of the model inference process of the target layer network structure usually has no effect. Therefore, the intermediate data of the model inference process of the target layer network structure can be deleted in the video card memory of the target deep learning model, and the corresponding video card memory can be released to reduce the video memory occupation during the model calculation process (for example, when the model inference of the target layer network structure ends, immediately clean the intermediate data (such as activation values) generated by it to free up more video memory space for the next layer network structure). In addition, the parallel feature of the video card can also be utilized to execute the step of deleting the intermediate data of the model inference process of the target layer network structure while starting the model inference of the next layer network structure.
[0073] In some embodiments of the present application, as Figure 3 shown, on the basis of the embodiments shown in Figure 1 or Figure 2 shown, after allocating video card memory for the target deep learning model according to the expected memory capacity, it may further include:
[0074] 301. In the model inference stage of the target deep learning model, decompose the matrix multiplication operation into multiple sub-operations.
[0075] In an embodiment of the present application, since some matrix multiplication operations are relatively complex and require a large amount of video card memory, during the model inference stage of the target deep learning model, the matrix multiplication operation can be decomposed into multiple sub-operations to simplify the complexity during model calculation.
[0076] 302. Use the video card memory of the target deep learning model to sequentially execute multiple sub-operations.
[0077] In an embodiment of the present application, using the video card memory of the target deep learning model to sequentially execute multiple sub-operations consumes less video card memory compared to directly executing the matrix multiplication operation.
[0078] In some embodiments of the present application, in order to obtain the result of the matrix multiplication operation, the memory allocation method for the deep learning model may further include: after each sub-operation is executed, sequentially accumulate the results of the executed sub-operations to obtain the result of the matrix multiplication operation, that is, after each sub-operation is executed, directly accumulate the result of the sub-operation, rather than waiting for all sub-operations to be executed and then merging the results of all sub-operations. Therefore, the video memory pressure can be further reduced. For example, a large matrix multiplication operation can be decomposed into several small-scale sub-operations to avoid occupying too much video memory at one time (for example, in the Attention mechanism, the projection calculations of the Query (query vector), Key (key vector), and Value (value vector) matrices can be completed in multiple small batches of sub-operations). And the result of each sub-operation is directly accumulated into the final output, rather than being all saved and then merged, further reducing the video memory pressure.
[0079] In some embodiments of the present application, as Figure 4 shown, on the basis of any one of the embodiments shown in Figures 1 to 3 after determining the expected memory capacity of the target deep learning model based on the number of model parameters, input size, and expected maximum batch size, it may further include:
[0080] 401. If the difference between the remaining capacity of the video card memory and the expected memory capacity is less than a preset difference, then reduce the expected maximum batch size according to the remaining capacity of the video card memory to adjust the expected memory capacity.
[0081] In an embodiment of the present application, if the difference between the remaining capacity of the video card memory and the expected memory capacity is less than a preset difference, it indicates that if the video card memory is directly allocated to the target deep learning model according to the expected memory capacity, phenomena such as memory overflow or a decrease in model inference speed may occur. Therefore, the expected memory capacity can be adjusted by reducing the expected memory capacity according to the remaining capacity of the video card memory to reduce the video card capacity required by the target deep learning model.
[0082] In some embodiments of the present application, reducing the expected maximum batch size according to the remaining capacity of the graphics card memory to adjust the expected memory capacity may include: reducing the expected maximum batch size by a preset amplitude so that the expected memory capacity is also reduced until the difference between the remaining capacity of the graphics card memory and the reduced expected memory capacity is greater than or equal to a preset difference.
[0083] When the difference between the remaining capacity of the graphics card memory and the reduced expected memory capacity is greater than or equal to the preset difference, it indicates that the current expected memory capacity and the expected maximum batch size (i.e., the reduced expected memory capacity and the expected maximum batch size) are appropriate. Therefore, the graphics card memory can be allocated to the target deep learning model according to the reduced expected memory capacity, and during the model inference stage of the target deep learning model, the target deep learning model can be inferred according to the reduced expected maximum batch size (i.e., for a target deep learning model that supports batch input, the entire expected maximum batch size can be split into smaller sub-batches for inference, which can reduce the demand for video memory in a single model inference), ensuring that there will be no phenomena such as out-of-memory or a decrease in the model inference speed during the actual process of model inference of the target deep learning model. In addition, during the process of inferring the target deep learning model according to the reduced expected maximum batch size, the corresponding results can be output in a timely manner after the inference of each reduced expected maximum batch size is completed, enabling the user to receive the corresponding results in a shorter time and also reducing the overall video memory burden on the host system.
[0084] In a second aspect, based on the above-mentioned memory allocation method for a deep learning model in the above embodiments, an embodiment of the present application provides a memory allocation device for a deep learning model. The memory allocation device for a deep learning model is used to execute the steps in any one of the above embodiments of the memory allocation method for a deep learning model. Specifically, the memory allocation device for a deep learning model may include:
[0085] A first acquisition module, configured to acquire the number of model parameters and the input size in the target deep learning model;
[0086] A second acquisition module, configured to acquire the expected maximum batch size of the target deep learning model;
[0087] A memory determination module, configured to determine the expected memory capacity of the target deep learning model based on the number of model parameters, the input size, and the expected maximum batch size;
[0088] A memory allocation module, configured to allocate the graphics card memory to the target deep learning model according to the expected memory capacity.
[0089] In a third aspect, embodiments of the present application provide a computer device that integrates any of the memory allocation devices for deep learning models provided in the embodiments of the present application. The computer device includes a processor and a memory. A computer program is stored in the memory and is configured to be executed by the processor to implement the memory allocation method for deep learning models described in any of the above embodiments. For example:
[0090] Obtain the number of model parameters and input size in the target deep learning model; obtain the expected maximum batch size of the target deep learning model; determine the expected memory capacity of the target deep learning model based on the number of model parameters, input size, and expected maximum batch size; and allocate video card memory for the target deep learning model according to the expected memory capacity.
[0091] In a fourth aspect, embodiments of the present application provide a computer device that integrates any of the memory allocation devices for deep learning models provided in the embodiments of the present application. As Figure 5 shown, it shows a schematic structural diagram of the computer device involved in the embodiments of the present application. Specifically:
[0092] The computer device may include a processor 501 with one or more processing cores, a storage unit 502 of one or more computer-readable storage media, a power supply 503, an input unit 504, and other components. Those skilled in the art can understand that Figure 5 the computer device structure shown in
[0093] does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0094] The storage unit 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the storage unit 502. The storage unit 502 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the storage unit 502 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the storage unit 502 may also include a memory controller to provide the processor 501 with access to the storage unit 502.
[0095] The computer device further includes a power supply 503 for supplying power to each component. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 503 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0096] The computer device may further include an input unit 504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0097] Although not shown, the computer device may further include a display unit, etc., which will not be elaborated here. Specifically, in the embodiment of the present application, the processor 501 in the computer device will load the executable files corresponding to the processes of one or more application programs into the storage unit 502 according to the following instructions, and the processor 501 will run the application programs stored in the storage unit 502 to realize various functions, such as:
[0098] Obtain the number of model parameters and input size in the target deep learning model; obtain the expected maximum batch processing volume of the target deep learning model; determine the expected memory capacity of the target deep learning model based on the number of model parameters, input size, and expected maximum batch processing volume; allocate graphics card memory for the target deep learning model according to the expected memory capacity.
[0099] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which may include: Read Only Memory (ROM), Random Access Memory (RAM), a magnetic disk, an optical disc, etc. The computer-readable storage medium stores a computer program, and the computer program is configured to be executed by a processor to implement the memory allocation method for a deep learning model as described in any one of the above, for example:
[0100] Obtain the number of model parameters and the input size in the target deep learning model; obtain the expected maximum batch processing amount of the target deep learning model; determine the expected memory capacity of the target deep learning model based on the number of model parameters, the input size, and the expected maximum batch processing amount; and allocate video card memory for the target deep learning model according to the expected memory capacity.
[0101] In a sixth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the memory allocation method for a deep learning model as described in any one of the above, for example:
[0102] Obtain the number of model parameters and the input size in the target deep learning model; obtain the expected maximum batch processing amount of the target deep learning model; determine the expected memory capacity of the target deep learning model based on the number of model parameters, the input size, and the expected maximum batch processing amount; and allocate video card memory for the target deep learning model according to the expected memory capacity.
[0103] The above has introduced the embodiments of the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A memory allocation method for a deep learning model, characterized in that The memory allocation method for the deep learning model includes: Obtain the number of model parameters and the input size in the target deep learning model; Obtain the expected maximum batch size of the target deep learning model; Based on the number of model parameters, the input size, and the expected maximum batch size, determine the expected memory capacity of the target deep learning model; Allocate video card memory for the target deep learning model according to the expected memory capacity.
2. The memory allocation method for a deep learning model according to claim 1, wherein After obtaining the number of model parameters in the target deep learning model, it further includes: Divide the model parameters in the target deep learning model into multiple parameter sets; After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes: In the model inference stage of the target deep learning model, determine the requested model parameters; Store the parameter set where the requested model parameters are located into the video card memory of the target deep learning model.
3. The memory allocation method for a deep learning model according to claim 2, wherein After storing the parameter set where the requested model parameters are located into the video card memory of the target deep learning model, it further includes: If the requested model parameters have been used up, determine the parameter type of the used-up model parameters; If the parameter type is a preset parameter type, transfer the used-up model parameters to the memory of the host where the video card is located for storage.
4. The memory allocation method for a deep learning model according to claim 1, wherein The target deep learning model includes a multi-layer network structure connected in sequence. After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes: In the multi-layer network structure connected in sequence, determine at least two network structures with the same weight matrix; In the video card memory of the target deep learning model, simultaneously allocate the memory area storing the same weight matrix to the at least two network structures.
5. The memory allocation method for a deep learning model according to claim 1, wherein The target deep learning model includes a multi-layer network structure connected in sequence. After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes: In the model inference stage of the target deep learning model, obtain the model inference result of the target layer network structure where the model inference has ended, so as to use the model inference result for the model inference of the next layer network structure of the target layer network structure; In the video card memory of the target deep learning model, delete the intermediate data of the model inference process of the target layer network structure, and the intermediate data is used to determine the model inference result.
6. The memory allocation method for a deep learning model according to claim 1, wherein, At least one layer network structure in the target deep learning model includes matrix multiplication operations. After allocating video card memory for the target deep learning model according to the expected memory capacity, it further includes: In the model inference stage of the target deep learning model, decompose the matrix multiplication operation into multiple sub-operations; Use the video card memory of the target deep learning model to sequentially execute the multiple sub-operations.
7. The memory allocation method for a deep learning model according to claim 6, wherein, The method further includes: After each sub-operation is executed, perform sequential accumulation processing on the result of the executed sub-operation to obtain the result of the matrix multiplication operation.
8. The memory allocation method for a deep learning model according to claim 1, wherein, After determining the expected memory capacity of the target deep learning model based on the number of model parameters, the input size, and the expected maximum batch size, the following steps are further included: If the difference between the remaining capacity of the graphics card memory and the expected memory capacity is less than a preset difference, then reduce the expected maximum batch size according to the remaining capacity of the graphics card memory to adjust the expected memory capacity; After allocating the graphics card memory to the target deep learning model according to the expected memory capacity, the following steps are further included: Perform model inference on the target deep learning model according to the reduced expected maximum batch size.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is configured to be executed by the processor to implement the memory allocation method for a deep learning model according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to be executed by a processor to implement the memory allocation method for a deep learning model according to any one of claims 1 to 8.