Neural network operator memory allocation method and system, electronic equipment and storage medium
By sorting and memory optimization strategies based on graph traversal algorithms in neural network models, the problem of insufficient memory allocation is solved, and the utilization rate of memory and model inference speed are improved.
Patent Information
- Application Number
- CN202510270803.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art fails to fully consider memory allocation optimization in the operator splitting process of neural network models, resulting in inefficient memory access and affecting the model inference speed.
By sorting the calculation order in the neural network model based on the preset graph traversal algorithm, memory allocation is performed in memory in turn, and splitting and memory multiplexing is performed when there is insufficient remaining space. Combined with the Ping-Pong memory allocation method, memory utilization is optimized.
It improves memory utilization, reduces the number of data handling times, and improves the inference speed of neural network models.
Smart Images

Figure CN120353572A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular, to a method and system for memory allocation of neural network operators, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of deep learning technology, the scale and complexity of neural network models have been continuously increasing, and the computational requirements of the models have grown exponentially. To meet these requirements, AI accelerators are widely used for accelerating deep learning tasks.
[0003] However, with the expansion of the model scale, large-grained operator operations (such as convolution, matrix multiplication, etc.) may consume more computational resources during the calculation process, and at the same time, they also increase the read and write requirements of memory and on-chip caches. For most AI accelerators, computational resources, memory bandwidth, and cache capacity are all precious resources, and how to efficiently utilize these resources has become a key challenge for improving the model inference speed.
[0004] Currently, most AI accelerators adopt a hierarchical storage structure, which usually includes external storage (such as DRAM) and internal storage (such as SRAM or on-chip cache). The external storage has a large capacity but a low bandwidth; the internal storage has a small capacity but a high bandwidth. Due to the limitation of the internal storage space, when an AI accelerator executes large-grained operators, it often needs to split the operators into multiple small-grained operations to adapt to the internal storage space.
[0005] The prior art describes how to split operators based on the hierarchical storage structure to ensure that computational tasks can be efficiently executed in the limited internal storage space. However, the inventors of the present invention have found that these operator splitting methods usually only focus on the splitting strategy of operators and ignore the optimization of memory allocation. That is, the inventors have found that the prior art fails to fully consider how to combine the optimized memory allocation when splitting operators, resulting in low memory access efficiency and affecting the model inference speed of neural networks. Summary of the Invention
[0006] The present invention provides a method and system for memory allocation of neural network operators, an electronic device, and a storage medium, aiming to solve the problem of low memory access efficiency existing in current operator splitting methods, which affects the model inference speed of neural networks.
[0007] According to one aspect of the present invention, the present invention provides a method for neural network operator memory allocation, including: sorting the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators; determining the target occupied memory required for the process data of the target operator; based on the target occupied memory of the target operator, sequentially performing memory allocation for the target operator in the first memory of the target chip; when the remaining storage space of the first memory meets the first preset condition, performing memory allocation for the currently to-be-memory-allocated target operator in the first memory; transmitting the current stored data in the first memory to the second memory of the target chip to continue traversing the memory allocation of the remaining target operators in the first memory in an idle state according to the calculation order until the memory allocation of all target operators is completed.
[0008] According to some embodiments of the present invention, performing memory allocation for the currently to-be-memory-allocated target operator in the first memory includes: splitting the currently to-be-memory-allocated target operator into multiple target sub-operators based on a preset splitting method; performing memory allocation for the multiple target sub-operators in the first memory based on a preset partition allocation method.
[0009] According to some embodiments of the present invention, based on the target occupied memory of the target operator, sequentially performing memory allocation for the target operator in the first memory of the target chip includes: when the operator adjacent to the currently to-be-memory-allocated target operator is a non-reusable operator, reusing the memory start address of the non-reusable operator as the start address of the currently to-be-memory-allocated target operator.
[0010] According to some embodiments of the present invention, the operator memory allocation method further includes: when the remaining storage space of the first memory meets the second preset condition, transmitting the current stored data in the first memory to the second memory of the target chip to traverse the memory allocation of the target operator in the first memory in an idle state until the memory allocation of all target operators is completed.
[0011] According to another aspect of the present invention, the present invention provides a neural network operator memory system, including an operator information processing unit and an operator memory allocation unit. The operator information processing unit sorts the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators, and determines the target occupied memory required for the process data of the target operators. The operator memory allocation unit performs memory allocation for the target operators in the first memory of the target chip in sequence based on the target occupied memory of the target operators. When the remaining storage space of the first memory meets the first preset condition, the operator memory allocation unit performs memory allocation for the currently to-be-memory-allocated target operator in the first memory, and transfers the current stored data in the first memory to the second memory of the target chip, so as to continue to traverse the memory allocation of the remaining target operators in the idle first memory according to the calculation order until the memory allocation of all target operators is completed.
[0012] According to some embodiments of the present invention, the operator memory allocation unit splits the currently to-be-memory-allocated target operator into multiple target sub-operators based on a preset splitting method. The operator memory allocation unit performs memory allocation for the multiple target sub-operators in the first memory based on a preset partition allocation method.
[0013] According to some embodiments of the present invention, when the operator adjacent to the currently to-be-memory-allocated target operator is a non-reusable operator, the operator memory allocation unit reuses the memory start address of the non-reusable operator as the start address of the currently to-be-memory-allocated target operator.
[0014] According to some embodiments of the present invention, when the remaining storage space of the first memory meets the second preset condition, the operator memory allocation unit transfers the current stored data in the first memory to the second memory of the target chip, so as to traverse the memory allocation of the target operators in the idle first memory until the memory allocation of all target operators is completed.
[0015] According to another aspect of the present invention, the present invention further provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, enable the one or more processors to implement the operator memory allocation method as described above.
[0016] According to another aspect of the present invention, the present invention further provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the operator memory allocation method as described above.
[0017] According to another aspect of the present invention, the present invention also provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions that, when executed by a computer, cause the computer to execute the operator memory allocation method as described above.
[0018] Beneficial effects
[0019] By sorting the network diagram structure in the neural network model of the target chip based on a preset graph traversal algorithm, the present invention can obtain the calculation order of all target operators, and realize memory allocation in the first memory according to the target occupied memory and the calculation order required by the process data of the target operators. In addition, when the remaining storage space of the first memory meets the first preset condition, the present invention can also perform memory allocation for the currently target operator to be memory-allocated in the first memory. Then, the current stored data in the first memory is transmitted to the second memory of the target chip, so as to continue traversing the memory allocation of the remaining target operators in the first memory in the idle state according to the calculation order until all target operator memory allocations are completed.
[0020] The present invention can sequentially store the target operators in the first memory based on the calculation order of the target operators. Combining with the memory optimization placement method, the memory utilization rate in the first memory can be improved. In addition, when the currently target operator to be memory-allocated cannot be accommodated by the remaining storage space of the first memory, the present invention can also perform memory allocation for the currently target operator to be memory-allocated in the first memory, which can reduce the number of data transfers to the second memory. The present invention can combine the optimized memory allocation to realize the splitting and memory allocation of operators, thereby improving the model inference speed of the neural network. Brief description of the drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 A flowchart showing the operator memory allocation method according to an embodiment of the present invention;
[0023] Figure 2 A structural diagram showing an AI accelerator according to an embodiment of the present invention;
[0024] Figure 3 A schematic diagram showing the operator memory allocation according to an embodiment of the present invention;
[0025] Figure 4 Another schematic diagram showing the operator memory allocation according to an embodiment of the present invention;
[0026] Figure 5 Another flowchart showing the operator memory allocation method according to an embodiment of the present invention;
[0027] Figure 6 Another schematic diagram showing the operator memory allocation according to an embodiment of the present invention;
[0028] Figure 7 Another flowchart showing the operator memory allocation method according to an embodiment of the present invention;
[0029] Figure 8 A schematic diagram showing the structure of the operator memory system according to an embodiment of the present invention.
[0030] Description of reference numerals:
[0031] Operator memory allocation system 1; Operator information processing unit 10; Operator memory allocation unit 20. Detailed implementation manners
[0032] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0033] The full names and Chinese interpretations of the English abbreviations involved in the present invention are as follows:
[0034] AI accelerator: Artificial Intelligence Accelerator, artificial intelligence accelerator;
[0035] GPU: Graphics Processing Unit, graphics processing unit;
[0036] TPU: Tensor Processing Unit, tensor processing unit;
[0037] FPGA: Field-Programmable Gate Array, field programmable gate array;
[0038] DRAM: Dynamic Random Access Memory, dynamic random access memory;
[0039] SRAM: Static Random Access Memory, static random access memory.
[0040] According to one aspect of the present invention, the present invention provides an operator memory allocation method for a neural network.
[0041] Figure 1 A flowchart showing the operator memory allocation method according to an embodiment of the present invention is shown. As shown in FIG. 1, the operator memory allocation method includes steps S100-S500. Exemplarily, the operator memory allocation method can be executed by an operator memory allocation system with computing capabilities.
[0042] According to an exemplary embodiment, in step S100, the operator memory allocation system sorts the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators.
[0043] For example, the target chip can be the chip for which operator memory allocation is to be performed. As an embodiment, the target chip can be an AI accelerator (including but not limited to GPU, TPU, and FPGA, etc.).
[0044] Figure 2 A structural diagram of an AI accelerator according to an embodiment of the present invention is shown.
[0045] As Figure 2 shown, the AI accelerator can include a microcontroller, an AI accelerator core, a first memory, and a second memory.
[0046] The microcontroller is the control core of the AI accelerator, responsible for coordinating and managing the work of each hardware module; the AI accelerator core is a computing unit specifically designed for AI tasks, used to execute computationally intensive tasks such as deep learning and machine learning; the first memory is the on-chip storage of the AI accelerator, such as SRAM; the second memory is the external storage of the AI accelerator, such as DRAM.
[0047] The network graph structure (or computational graph) in the neural network model of the target chip is a graphical representation for describing the calculation process of the neural network model, which includes nodes and edges. Nodes represent the computational operations in the neural network model, that is, operators; edges represent the flow direction of operators between nodes, that is, describe the input and output relationships of computational operations.
[0048] According to an exemplary embodiment, the preset graph traversal algorithm can be a depth-first algorithm.
[0049] The operator memory allocation system can sort the target operators (referring to the operators to be calculated in the network graph structure) of the network graph structure based on the depth-first algorithm to obtain the calculation order of all target operators in the network graph structure. With such a setting, the present invention enables the operator memory allocation system to sequentially perform operator memory allocation on all target operators based on this fixed calculation order.
[0050] In step S200, the operator memory allocation system determines the target occupied memory required for the process data of the target operator.
[0051] For example, the process data includes at least one of the input data of the target operator, the output data of the target operator, and the parameter data of the target operator. Exemplarily, the process data includes, but is not limited to, data forms such as image data, text data, audio-video data, or structured data.
[0052] The operator memory allocation system calculates the size of the memory space required for the input data, output data, or parameter data of each target operator to obtain the target occupied memory corresponding to each target operator.
[0053] According to the exemplary embodiment, the types of target operators may include basic mathematical operation operators (i.e., numerical calculation operators) and neural network layer operators. In the case where the target operator is a basic mathematical operation operator, the operator memory allocation system only needs to calculate one copy for its input data and output data.
[0054] It can be understood here that when calculating the occupied memory of the basic mathematical operation operator, only the memory space occupied by the input data and output data of a single sample needs to be considered, that is, the memory occupation of a single sample can represent the memory occupation mode of the entire operator. With such a setting, the present invention does not need to perform repeated calculations, thereby simplifying the memory estimation process.
[0055] In step S300, the operator memory allocation system sequentially allocates memory for the target operator in the first memory of the target chip based on the target occupied memory of the target operator.
[0056] For example, the calculation order of the target operators may be: the first target operator, the second target operator... the Nth target operator, where N is the total number of target operators.
[0057] The operator memory allocation system sequentially combines the first target operator, the second target operator... the Nth target operator with their corresponding target occupied memories and allocates corresponding memory for them in the first memory.
[0058] Figure 3 A schematic diagram showing the operator memory allocation according to an embodiment of the present invention; Figure 4 Another schematic diagram showing the operator memory allocation according to an embodiment of the present invention; Figure 5 Another flowchart schematic diagram showing the operator memory allocation method according to an embodiment of the present invention; Figure 6 Another schematic diagram showing the operator memory allocation according to an embodiment of the present invention.
[0059] As an embodiment, as Figure 3As shown, the operator memory allocation system sequentially allocates corresponding memory for the first target operator 1, the second target operator 2, the third target operator 3, and the fourth target operator 4 in the first memory in order.
[0060] Optionally, in step S300, when the operator adjacent to the target operator currently waiting for memory allocation is a non-reusable operator, the operator memory allocation system reuses the memory start address of the non-reusable operator as the start address of the target operator currently waiting for memory allocation.
[0061] For example, if the output data of a certain target operator is no longer used in subsequent calculations, the operator memory allocation system can mark this target operator as a non-reusable operator. During the memory allocation process of the target operator, if the operator adjacent to (or approximately adjacent to) the target operator currently waiting for memory allocation is a non-reusable operator (or can also be called a discarded operator), the operator memory allocation system can use the start address of this non-reusable operator as the start address of the target operator currently waiting for memory allocation to perform corresponding memory allocation.
[0062] As an embodiment, as Figure 4 shown, when the third target operator 3 is a non-reusable operator, the fifth target operator can reuse the start address of the third target operator 3.
[0063] Through the above embodiments, in order to efficiently manage the memory allocation in the first memory, the present invention adopts an optimized memory allocation strategy, combines the memory reuse and discarded memory recycling mechanisms, marks the memory occupied by the discarded operator as discarded memory for subsequent use. With such a setting, the present invention can reduce memory fragmentation and improve the memory utilization rate of on-chip storage resources.
[0064] In step S400, when the remaining storage space in the first memory meets the first preset condition, the operator memory allocation system performs memory allocation for the target operator currently waiting for memory allocation in the first memory.
[0065] For example, the first preset condition can be: the remaining storage space in the first memory is less than the target occupied space of the target operator currently waiting for memory allocation, and the remaining storage space in the first memory is greater than or equal to a preset threshold. Denote the minimum amount of computation required when the target operator currently waiting for memory allocation is calculated in the AI accelerator core as M, then the preset threshold is 2M.
[0066] Optionally, as Figure 5 shown, step S400 may further include steps S410 - S420.
[0067] In step S410, the operator memory allocation system splits the target operator currently waiting for memory allocation into multiple target sub-operators based on a preset splitting method.
[0068] For example, the preset splitting method may be to use the minimum computation amount M as the splitting unit to split the target occupied memory of the target operator to be currently allocated memory.
[0069] Exemplarily, when the remaining storage space of the first memory is less than the target occupied space of the target operator to be currently allocated memory and greater than or equal to 2M, the operator memory allocation system splits the target operator to be currently allocated memory with M as the splitting unit, so as to obtain N target sub-operators S1 (that is, the size of the occupied memory required for each target sub-operator S1 is M).
[0070] In step S420, the operator memory allocation system performs memory allocation for multiple target sub-operators in the first memory based on a preset partition allocation method.
[0071] Optionally, the preset partition allocation method is the Ping-Pong memory allocation method.
[0072] The Ping-Pong memory allocation method (Ping-Pong Mode) can perform data transfer and task calculation through a partition mechanism (double buffer, including a Ping area and a Pong area), and data transfer and task calculation can be alternately performed in the Ping area and the Pong area.
[0073] For example, the operator memory allocation system allocates a total of 2M of memory space for N target sub-operators S1 in the first memory (that is, in the present invention, the sizes of the Ping area and the Pong area are both 1M.), and configures the starting address of the occupied memory for N target sub-operators S1 in the allocated 2M of memory space to perform corresponding memory allocation.
[0074] Through the above embodiments, the present invention enables multiple target sub-operators to perform memory allocation in a Ping-Pong manner in the memory space of the first memory, so that multiple target sub-operators can transfer and calculate at the same time. With such a setting, the present invention can reduce the number of times of data transfer to the second memory and improve the memory utilization rate of the AI accelerator.
[0075] In step S500, the operator memory allocation system transfers the stored data in the first memory to the second memory of the target chip, so as to continue to traverse the memory allocation of the remaining target operators according to the calculation order in the first memory in the idle state until the memory allocation of all target operators is completed.
[0076] For example, after multiple target sub-operators perform memory allocation based on the Ping-Pong memory allocation method, the AI accelerator kernel outputs the memory allocation result to the second memory and transfers the current stored data in the first memory to the second memory.
[0077] Exemplarily, the current stored data in the first memory is the stored data that has not been recycled. During subsequent operator calculations, if this data is still needed, the operator memory allocation system can also transfer the stored data back to the previous storage location in the first memory to facilitate subsequent operator calculations.
[0078] After the operator memory allocation system transfers the stored data in the first memory to the second memory, the first memory becomes idle again. Based on the above operator memory allocation method, the operator memory allocation system continues to traverse the remaining target operators based on the calculation order of the target operators until the remaining storage space in the first memory again meets the first preset condition. Then, the operator memory allocation system splits the target operator again and performs memory allocation for multiple target sub-operators in the first memory again in the preset partition allocation manner. This is executed in a loop until the memory allocation for all target operators is completed.
[0079] According to the exemplary embodiment, the operator memory allocation system can group the target operators that perform memory allocation in the same batch in the first memory.
[0080] Exemplarily, as Figure 6 shown, the target operators 1-5 that perform operator memory allocation in the first memory in the first batch are Group1; the target operators 6-10 that perform operator memory allocation in the first memory in the second batch are Group2, and so on.
[0081] Figure 7 Another flowchart showing the operator memory allocation method according to an embodiment of the present invention is shown. As Figure 7 shown, the operator memory allocation method may further include step S600.
[0082] Optionally, in step S600, when the remaining storage space in the first memory meets the second preset condition, the operator memory allocation system transfers the current stored data in the first memory to the second memory of the target chip to traverse the memory allocation of the target operators in the idle first memory until the memory allocation for all target operators is completed.
[0083] For example, the second preset condition may be: the remaining storage space of the first memory is less than the target occupied space of the target operator to be currently memory - allocated, and the remaining storage space of the first memory is less than a preset threshold. Denote the minimum amount of computation required when the target operator to be currently memory - allocated is computed in the AI accelerator kernel as M, then the preset threshold is 2M.
[0084] Exemplarily, when the remaining storage space of the first memory is less than the target occupied space of the target operator to be currently memory - allocated and less than 2M, the operator memory allocation system transfers the current stored data in the first memory to the second memory so that the first memory becomes idle again. The operator memory allocation system then continues to perform the memory allocation for the next batch of operators (i.e., the target operators that meet the second preset condition enter the memory allocation calculation of the next Group). This is executed in a loop until the memory allocation for all target operators is completed.
[0085] Through the above - mentioned embodiments, the present invention can sort the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators, and can achieve memory allocation in the first memory according to the target occupied memory and the calculation order of the process data of the target operators. And the present invention can also, when the remaining storage space of the first memory meets the first preset condition, split the target operator to be currently memory - allocated into multiple target sub - operators based on a preset splitting method, and perform memory allocation for the multiple target sub - operators in the first memory based on a preset partition allocation method. Then, the current stored data in the first memory is transmitted to the second memory of the target chip to continue traversing the memory allocation of the remaining target operators in the idle first memory according to the calculation order until the memory allocation for all target operators is completed.
[0086] The present invention can store the target operators in the first memory in sequence according to the calculation order of the target operators. Combined with the memory - optimized placement method, it can improve the memory utilization rate in the first memory. And the present invention can also, when the target operator to be currently memory - allocated cannot be accommodated by the remaining storage space of the first memory, split the target operator to be currently memory - allocated into multiple target sub - operators based on the minimum amount of computation required when it is computed in the AI accelerator kernel, and perform memory allocation for the multiple target sub - operators in the first memory based on a preset partition allocation method, which can reduce the number of times of data transfer to the second memory. The present invention can combine the optimized memory allocation to achieve the splitting and memory allocation of operators, thereby improving the model inference speed of the neural network.
[0087] According to another aspect of the present invention, the present invention also provides a neural network operator memory allocation system.
[0088] Figure 8 A schematic structural diagram of the operator memory system according to an embodiment of the present invention is shown. As Figure 8 shown, the operator memory allocation system 1 may include an operator information processing unit 10 and an operator memory allocation unit 20.
[0089] According to an exemplary embodiment, the operator information processing unit 10 sorts the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators.
[0090] For example, the target chip may be the chip for which operator memory allocation is to be performed. As an embodiment, the target chip may be an AI accelerator (including but not limited to GPU, TPU, FPGA, etc.).
[0091] As Figure 2 shown, the AI accelerator may include a microcontroller, an AI accelerator core, a first memory, and a second memory.
[0092] The microcontroller is the control core of the AI accelerator, responsible for coordinating and managing the work of each hardware module; the AI accelerator core is a computing unit specially designed for AI tasks, used to execute computationally intensive tasks such as deep learning and machine learning; the first memory is the on-chip storage of the AI accelerator, such as SRAM; the second memory is the external storage of the AI accelerator, such as DRAM.
[0093] The network graph structure (or computational graph) in the neural network model of the target chip is a graphical representation for describing the calculation process of the neural network model, which includes nodes and edges. The nodes represent the computational operations in the neural network model, that is, operators; the edges represent the flow direction of the operators between the nodes, that is, describe the input and output relationships of the computational operations.
[0094] According to an exemplary embodiment, the preset graph traversal algorithm may be a depth-first algorithm.
[0095] The operator information processing unit 10 may sort the target operators (referring to the operators to be calculated in the network graph structure) of the network graph structure based on the depth-first algorithm to obtain the calculation order of all target operators in the network graph structure. With this setting, the present invention enables the operator memory allocation system to sequentially perform operator memory allocation on all target operators based on this fixed calculation order.
[0096] The operator information processing unit 10 determines the target occupied memory required for the process data of the target operator.
[0097] For example, the process data includes at least one of the input data of the target operator, the output data of the target operator, and the parameter data of the target operator. Exemplarily, the process data includes, but is not limited to, data forms such as image data, text data, audio-visual data, or structured data.
[0098] The operator information processing unit 10 calculates the size of the memory space required for the input data, output data, or parameter data of each target operator to obtain the target occupied memory corresponding to each target operator.
[0099] According to the exemplary embodiment, the types of target operators may include basic mathematical operation operators (i.e., numerical calculation operators) and neural network layer operators. In the case where the target operator is a basic mathematical operation operator, the operator information processing unit 10 only needs to calculate one copy of its input data and output data.
[0100] It can be understood here that when calculating the occupied memory of the basic mathematical operation operator, only the memory space occupied by the input data and output data of a single sample needs to be considered, that is, the memory occupation of a single sample can represent the memory occupation mode of the entire operator. With such a setting, the present invention does not need to perform repeated calculations, thereby simplifying the memory estimation process.
[0101] The operator memory allocation unit 20 allocates memory for the target operator in the first memory of the target chip in sequence based on the target occupied memory of the target operator.
[0102] For example, the calculation order of the target operators may be: the first target operator, the second target operator... the Nth target operator, where N is the total number of target operators.
[0103] The operator memory allocation system sequentially combines the first target operator, the second target operator... the Nth target operator with their corresponding target occupied memory and allocates corresponding memory for them in the first memory.
[0104] As an embodiment, as Figure 3 shown, the operator memory allocation system sequentially allocates corresponding memory for the first target operator 1, the second target operator 2, the third target operator 3, and the fourth target operator 4 in the first memory.
[0105] Optionally, in the case where the operator adjacent to the currently to-be memory-allocated target operator is a non-reusable operator, the operator memory allocation unit 20 reuses the memory start address of the non-reusable operator as the start address of the currently to-be memory-allocated target operator.
[0106] For example, if the output data of a certain target operator is no longer used in subsequent calculations, the operator memory allocation unit 20 can mark the target operator as a non-reusable operator. During the memory allocation process of the target operator, if an adjacent operator (or approximately adjacent operator) to the currently memory-allocation target operator is a non-reusable operator (or can also be called a discarded operator), the operator memory allocation unit 20 can use the starting address of the non-reusable operator as the starting address of the currently memory-allocation target operator to perform corresponding memory allocation.
[0107] As an embodiment, as Figure 4 shown, in the case where the third target operator 3 is a non-reusable operator, the fifth target operator can reuse the starting address of the third target operator 3.
[0108] Through the above embodiments, in order to efficiently manage the memory allocation in the first memory, the present invention adopts an optimized memory allocation strategy, combines the memory reuse and discarded memory recycling mechanisms, marks the memory occupied by the discarded operator as discarded memory for subsequent use. With such a setting, the present invention can reduce memory fragmentation and improve the memory utilization rate of the on-chip storage resources.
[0109] According to the exemplary embodiment, the operator memory allocation unit 20 performs memory allocation for the currently memory-allocation target operator in the first memory when the remaining storage space of the first memory meets the first preset condition.
[0110] For example, the first preset condition can be: the remaining storage space of the first memory is less than the target occupied space of the currently memory-allocation target operator, and the remaining storage space of the first memory is greater than or equal to a preset threshold. Denote the minimum amount of computation required when the currently memory-allocation target operator is calculated in the AI accelerator core as M, then the preset threshold is 2M.
[0111] Optionally, the operator memory allocation unit 20 splits the currently memory-allocation target operator into multiple target sub-operators based on a preset splitting method.
[0112] For example, the preset splitting method can be to split the target occupied memory of the currently memory-allocation target operator with M as the splitting unit.
[0113] Exemplarily, when the remaining storage space of the first memory is less than the target occupied space of the currently memory-allocation target operator and greater than or equal to 2M, the operator memory allocation unit 20 splits the currently memory-allocation target operator with M as the splitting unit, thereby obtaining N target sub-operators S1 (that is, the size of the occupied memory required for each target sub-operator S1 is M).
[0114] The operator memory allocation unit 20 performs memory allocation for multiple target sub-operators in the first memory based on a preset partition allocation method.
[0115] Optionally, the preset partition allocation method is the Ping-Pong memory allocation method.
[0116] The Ping-Pong memory allocation method (Ping-Pong Mode) can perform data transfer and task calculation through a partitioning mechanism (double buffer, including a Ping area and a Pong area), and data transfer and task calculation can be alternated between the Ping area and the Pong area.
[0117] For example, the operator memory allocation unit 20 allocates a total of 2M of memory space for N target sub-operators S1 in the first memory (that is, in the present invention, the sizes of both the Ping area and the Pong area are 1M.), and configures the starting address of the occupied memory for the N target sub-operators S1 in the allocated 2M of memory space based on the Ping-Pong memory allocation method to perform corresponding memory allocation.
[0118] Through the above embodiments, the present invention enables multiple target sub-operators to perform memory allocation in a Ping-Pong manner in the memory space of the first memory, so that multiple target sub-operators can transfer and calculate at the same time. With such a setting, the present invention can reduce the number of times of data transfer to the second memory and improve the memory utilization rate of the AI accelerator.
[0119] The operator memory allocation unit 20 transfers the stored data in the first memory to the second memory of the target chip, so as to continue to traverse the memory allocation of the remaining target operators according to the calculation order in the first memory in an idle state until the memory allocation of all target operators is completed.
[0120] For example, after the memory allocation of multiple target sub-operators is performed based on the Ping-Pong memory allocation method, the AI accelerator kernel outputs the memory allocation result to the second memory and transfers the current stored data in the first memory to the second memory.
[0121] Exemplarily, the current stored data in the first memory is the stored data that has not been recycled. In the subsequent operator calculation process, if the data is still needed, the operator memory allocation unit 20 can also transfer the stored data back to the previous storage location in the first memory for subsequent operator calculation.
[0122] After the operator memory allocation unit 20 transfers the stored data in the first memory to the second memory, the first memory returns to the idle state again. Based on the above operator memory allocation method, the operator memory allocation unit 20 continues to traverse the remaining target operators according to the calculation order of the target operators until the remaining storage space of the first memory satisfies the first preset condition again. Then, the operator memory allocation unit 20 splits the target operator again and performs the memory allocation of multiple target sub-operators in the first memory again in the preset partition allocation manner. This process is executed in a loop until the memory allocation of all target operators is completed.
[0123] According to the exemplary embodiment, the operator memory allocation unit 20 can group the target operators for which memory allocation is performed in the same batch in the first memory.
[0124] Exemplarily, as Figure 5 shown, the target operators 1-5 for which operator memory allocation is performed in the first memory in the first batch are Group1; the target operators 6-10 for which operator memory allocation is performed in the first memory in the second batch are Group2, and so on.
[0125] Optionally, when the remaining storage space of the first memory satisfies the second preset condition, the operator memory allocation unit 20 transfers the current stored data in the first memory to the second memory of the target chip to traverse the memory allocation of the target operators in the idle first memory until the memory allocation of all target operators is completed.
[0126] For example, the second preset condition can be: the remaining storage space of the first memory is less than the target occupied space of the current target operator to be memory-allocated, and the remaining storage space of the first memory is less than the preset threshold. Denote the minimum amount of computation required when the current target operator to be memory-allocated is computed in the AI accelerator core as M, then the preset threshold is 2M.
[0127] Exemplarily, when the remaining storage space of the first memory is less than the target occupied space of the current target operator to be memory-allocated and less than 2M, the operator memory allocation unit 20 transfers the current stored data in the first memory to the second memory so that the first memory returns to the idle state again. The operator memory allocation unit 20 then continues to perform the memory allocation of the next batch (i.e., the target operators that meet the second preset condition enter the memory allocation calculation of the next Group). This process is executed in a loop until the memory allocation of all target operators is completed.
[0128] Through the above embodiments, the present invention can sort the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators. Memory allocation in the first memory can be achieved according to the target occupied memory and the calculation order required by the process data of the target operators. In addition, when the remaining storage space of the first memory meets the first preset condition, the present invention can split the target operator to be memory-allocated currently into multiple target sub-operators based on a preset splitting method, and perform memory allocation for the multiple target sub-operators in the first memory based on a preset partition allocation method. Then, the current stored data in the first memory is transmitted to the second memory of the target chip, so as to continue traversing the memory allocation of the remaining target operators according to the calculation order in the first memory in an idle state until the memory allocation of all target operators is completed.
[0129] The present invention can store the target operators in the first memory in sequence according to the calculation order of the target operators. Combined with the memory optimization placement method, the memory utilization rate in the first memory can be improved. In addition, when the target operator to be memory-allocated currently cannot be accommodated by the remaining storage space of the first memory, the present invention can split the target operator to be memory-allocated currently into multiple target sub-operators based on the minimum calculation amount required during its calculation in the AI accelerator kernel, and perform memory allocation for the multiple target sub-operators in the first memory based on a preset partition allocation method, which can reduce the number of times of data transfer to the second memory. The present invention can combine the optimized memory allocation to realize the splitting and memory allocation of operators, thereby improving the model inference speed of the neural network.
[0130] According to another aspect of the present invention, there is also provided an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the operator memory allocation method as described above.
[0131] According to another aspect of the present invention, there is also provided a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the operator memory allocation method as described above.
[0132] According to another aspect of the present invention, there is also provided a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the operator memory allocation method as described above.
[0133] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions of the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for allocating memory of a neural network operator, characterized in that including: sorting the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators; determining the target memory occupied by the process data of the target operator; based on the target memory occupied by the target operator, sequentially performing memory allocation for the target operator in the first memory of the target chip according to the calculation order; when the remaining storage space in the first memory meets the first preset condition, performing memory allocation for the target operator to be currently memory-allocated in the first memory; transmitting the current stored data in the first memory to the second memory of the target chip, so as to continue traversing the memory allocation of the remaining target operators according to the calculation order in the first memory in an idle state until the memory allocation of all the target operators is completed.
2. The operator memory allocation method according to claim 1, wherein The performing memory allocation for the target operator to be currently memory-allocated in the first memory includes: splitting the target operator to be currently memory-allocated into multiple target sub-operators based on a preset splitting method; performing memory allocation for the multiple target sub-operators in the first memory based on a preset partition allocation method.
3. The operator memory allocation method according to claim 1, wherein The sequentially performing memory allocation for the target operator in the first memory of the target chip based on the target memory occupied by the target operator according to the calculation order includes: when the operator adjacent to the target operator to be currently memory-allocated is a non-reusable operator, reusing the memory start address of the non-reusable operator as the start address of the target operator to be currently memory-allocated.
4. The operator memory allocation method according to claim 1, wherein The operator memory allocation method further includes: when the remaining storage space in the first memory meets the second preset condition, transmitting the current stored data in the first memory to the second memory of the target chip, so as to traverse the memory allocation of the target operator in the first memory in an idle state until the memory allocation of all the target operators is completed.
5. A neural network operator memory allocation system, characterized in that, including: an operator information processing unit, sorting the network graph structure in the neural network model of the target chip based on a preset graph traversal algorithm to obtain the calculation order of all target operators, and determining the target memory occupied by the process data of the target operator; an operator memory allocation unit, sequentially performing memory allocation for the target operator in the first memory of the target chip based on the target memory occupied by the target operator according to the calculation order, when the remaining storage space in the first memory meets the first preset condition, performing memory allocation for the target operator to be currently memory-allocated in the first memory, and transmitting the current stored data in the first memory to the second memory of the target chip, so as to continue traversing the memory allocation of the remaining target operators according to the calculation order in the first memory in an idle state until the memory allocation of all the target operators is completed.
6. The operator memory allocation system according to claim 5, wherein The operator memory allocation unit splits the target operator to be currently memory-allocated into multiple target sub-operators based on a preset splitting method; The operator memory allocation unit performs memory allocation for the multiple target sub-operators in the first memory based on a preset partitioning and allocation method.
7. The operator memory allocation system according to claim 5, wherein When the operator adjacent to the target operator currently to be allocated memory is a non-reusable operator, the operator memory allocation unit reuses the memory start address of the non-reusable operator as the start address of the target operator currently to be allocated memory.
8. The operator memory allocation system according to claim 5, wherein When the remaining storage space in the first memory meets a second preset condition, the operator memory allocation unit transfers the currently stored data in the first memory to the second memory of the target chip, so as to traverse the memory allocation of the target operators in the first memory in an idle state until the memory allocation of all the target operators is completed.
9. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the operator memory allocation method according to any one of claims 1-4.
10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the operator memory allocation method according to any one of claims 1-4.