Efficient memory reuse method for edge device storage feature map

By dividing deep neural network nodes into ordinary nodes and splicing nodes, and using the best matching and loose matching search strategies for memory address allocation, the problems of insufficient memory space and high power consumption of edge devices are solved, and more efficient memory utilization and computing performance are achieved.

CN119948462APending Publication Date: 2025-05-06HONG KONG APPLIED SCI & TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480003263.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-04
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When running deep neural network models, edge devices face problems such as insufficient memory space and high power consumption, especially when multiple image processing AI models are needed to run simultaneously.

Method used

A memory reuse algorithm is proposed to divide deep neural network nodes into ordinary nodes and splicing nodes. Through the best matching search strategy and loose matching search strategy, the allocation of memory addresses is optimized, the efficient use of memory is ensured, and power consumption is reduced.

Benefits of technology

Through this algorithm, edge devices can more efficiently utilize memory resources, reduce memory access latency, improve computing efficiency and response speed, support the operation of more complex neural network models, and reduce power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948462A_ABST
    Figure CN119948462A_ABST
Patent Text Reader

Abstract

A system and method optimizes memory utilization when executing neural networks (DNNs) on devices such as resource-constrained edge devices. The method comprises the following steps: dividing neural network nodes into common nodes and splicing nodes; sorting the common nodes according to a descending order of the sizes of the feature graphs; allocating memory addresses for the common nodes by using optimal matching search; and allocating a memory address for the splicing node using a loose match search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for efficiently utilizing memory space in a deep neural network model, which can be used to store intermediate feature maps. Background Art

[0002] The current development of edge devices and neural processing units (NPUs) is receiving widespread attention and has become a highly discussed topic. Traditional edge devices, such as Tesla's cars, need to run 50 different image processing artificial intelligence (AI) models simultaneously. In addition, a typical artificial intelligence (AI) model in the field of image processing usually requires about 100MB to 200MB (@INT16) of memory space to store parameters and feature maps. A typical NPU for inference on an edge device is usually equipped with an on-chip memory (SRAM) of 512KB to 32MB. Edge devices usually have strict budget constraints on memory (DRAM), and the memory consumption of feature maps can be reduced through a memory reuse mechanism. In addition, some feature maps can also be stored in SRAM, thereby reducing power consumption. The present invention proposes a memory reuse algorithm that can efficiently allocate the memory address of artificial intelligence (AI) models in SRAM and DRAM, thereby saving memory space and power consumption of the NPU. Summary of the invention

[0003] In one aspect, the proposed method involves classifying deep neural network nodes into two categories: normal nodes and spliced ​​nodes. Then, the normal nodes are sorted in descending order according to the size of their feature maps. The memory allocation of normal nodes can be done by the best matching search strategy. For spliced ​​nodes, a more flexible loose matching search strategy is used to allocate memory addresses.

[0004] On the other hand, the proposed method improves the memory efficiency of neural network processing in edge devices by classifying nodes according to their memory requirements and using different memory allocation algorithms for each category. The proposed method strategically uses SRAM and DRAM for memory allocation and optimizes memory addresses for optimal efficiency and reliability. In addition, the proposed method preferentially allocates nodes with shorter lifecycles and less memory required for feature maps to SRAM.

[0005] Furthermore, the edge device contains components specifically designed to manage neural network nodes. The device used has a mechanism to classify nodes into two categories: normal nodes and spliced ​​nodes. After the classification is completed, the device sorts the normal nodes in descending order according to the size of their feature graphs. In terms of memory allocation, the device adopts the best match search strategy for normal nodes to ensure efficient use of memory. For spliced ​​nodes, the device adopts a loose match search strategy to allocate memory addresses, thereby providing flexibility in memory management.

[0006] An embodiment of the present invention may include one or more of the following advantages:

[0007] One implementation improves the computational efficiency of neural networks on edge devices by reducing memory-related bottlenecks, ensuring that devices with limited computing power can still perform the complex calculations required for artificial intelligence (AI) applications.

[0008] By introducing best-fit and loose-fit allocation methods for different node types, one implementation maximizes the efficient use of available memory, allowing larger and potentially more complex neural networks to run on resource-constrained devices.

[0009] Memory allocation strategies help minimize the delays that are typically associated with memory access times. This is because the proposed approach aims to ensure that memory is used as efficiently as possible, thereby reducing delays caused by memory paging or excessive data swapping with external memory sources.

[0010] Sorting the feature maps of common nodes in descending order can enhance memory locality, thereby improving cache performance and further reducing access time, thereby providing faster and more responsive artificial intelligence (AI) systems on edge devices.

[0011] By pragmatically utilizing SRAM and DRAM in the allocation process, one embodiment can take advantage of the faster access speed of SRAM for more critical or frequently accessed nodes, while simultaneously taking advantage of the larger storage capacity of DRAM to store less critical data.

[0012] This customized approach of prioritizing SRAM allocation to nodes with shorter lifecycles meets the real-time processing requirements of edge computing applications, where timely data processing is critical.

[0013] Dynamic memory allocation technology can avoid memory fragmentation and effectively manage write operations to various types of memory (volatile and non-volatile), thereby reducing wear and tear to improve the reliability of edge devices and extend their service life.

[0014] The advanced memory management strategy advocated by this embodiment can provide a competitive advantage to devices operating in the Internet of Things, autonomous systems, and real-time analytics, where processing efficiency and fast decision making are essential.

[0015] This implementation facilitates the scalability of the neural network because efficient memory management can accommodate the growth of model complexity when new layers or nodes are added.

[0016] Finally, this innovative approach can save energy by reducing the need for external data storage solutions, which typically require additional power and may introduce communication overhead, resulting in more energy consumption.

[0017] The proposed technology adopts selective memory allocation strategies, including best match and loose match strategies, aiming to improve the efficiency of memory address allocation. In addition, the system can flexibly handle memory allocation of static (SRAM) and dynamic (DRAM) random access memories, impose constraints on the size and life cycle of nodes, and prioritize SRAM allocation to nodes with specific characteristics. This not only saves memory space, but also ensures that nodes are connected and operated in sequence, which is particularly important for edge devices with limited computing power and memory capacity. By implementing the system, complex deep neural network operations can be performed more efficiently on edge devices, enabling them to support advanced functions that are traditionally difficult to achieve due to power consumption and space limitations. Overall, the proposed implementation significantly improves the performance of neural networks in edge devices, laying the foundation for the ability to achieve advanced artificial intelligence (AI) in a world that increasingly relies on intelligent and autonomous systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 An example flow chart for allocating memory addresses of neural network nodes is presented.

[0019] Figure 2 An exemplary flowchart for efficient memory management of feature maps in neural networks deployed at the edge is presented.

[0020] Figure 3 A flowchart of the best match search strategy for efficient memory management in neural networks deployed at the edge is presented.

[0021] Figures 4 to 5 Example pseudocodes for static memory planning and best match search strategies are presented respectively.

[0022] FIG. 6A to FIG. 6B Example pseudocode and flow chart for the relaxed matching search strategy process are presented.

[0023] Figure 7 shows Figure 1 The table of exemplary operations in the process of Figure 1 The storage savings that can be achieved through the process. DETAILED DESCRIPTION

[0024] A method for efficiently utilizing memory space to store the intermediate feature graphs of a deep neural network model through a memory planning algorithm. The memory planning algorithm divides nodes into two categories: one is ordinary nodes, which are arranged in descending order of their feature graph sizes; the other is nodes that need to be spliced ​​together. The set of nodes that need to be spliced ​​together is called a concat tuple. The memory planning algorithm can use the best match search algorithm to allocate memory addresses for ordinary nodes, and use the loose match search algorithm to allocate memory addresses for spliced ​​nodes.

[0025] The memory planning algorithm allocates memory for the intermediate feature graph of each node in the static random access memory (SRAM) and the dynamic random access memory (DRAM), preferentially allocating memory in the SRAM and then allocating memory in the DRAM, and allocates according to the provided size constraints (i.e., the SRAM size constraint and the DRAM size constraint) and the life cycle constraints of the node allocation to the SRAM and DRAM, so that the SRAM can avoid processing nodes with too long a life cycle. In one example, the best match search algorithm may try the following: i. From the list of allocated nodes, obtain a list of nodes that overlap with the current node's life cycle and have been allocated addresses; ii. Find the list of memory spaces divided by allocated and overlapping nodes; iii. Find the minimum space that can accommodate the current node; iv. Assign the starting address of the space as the address of the current node; v. Insert the current node into the list of allocated nodes and sort the list of allocated nodes in ascending order of address. Loose matching search algorithm: i. Divide the concatenated tuple into two subtypes: I. List of concatenated tuples with unique nodes (subtype 1): These concatenated tuples have no shared / common nodes with other concatenated tuples; II. List of spliced ​​tuples with shared nodes (subtype 2): These spliced ​​tuples have shared / common nodes with other spliced ​​tuples; III. The concatenated tuple list of unique nodes can be allocated in SRAM or DRAM; IV. The concatenated tuple list of a shared node can only be allocated in DRAM. ii. For each tuple in the concatenated tuple list of unique nodes: I. Group all nodes in the same tuple into a supernode where: 1) The starting node is the minimum starting node of all nodes in the tuple; 2) The end node is the maximum end node of all nodes in the tuple; 3) The size is the sum of the feature map sizes of all nodes in the tuple. II. The supernode will obtain an address in memory through the best match search algorithm: 1) The address of the first node in the tuple is the address of the supernode; 2) The address of the next node ([n+1]) is the end address of the current node ([n]), that is, the address of the current node plus the output feature map size of the current node; 3) Insert all nodes in the same tuple into the allocated node list and sort them in ascending order of address. iii. For each tuple in the shared node concatenate tuple list: I. All nodes in the same tuple will first be assigned an address that follows the largest ending address of all overlapping nodes that have previously been assigned addresses. II. “All overlapping nodes with allocated addresses” do not include previously allocated shared nodes. III. Insert all nodes into the allocated node list and sort them in ascending order of address. IV. The address of each node will be adjusted to maintain its order in the concatenation operator: if the address of the next node is less than the end address of the current node, the address of the next node will be assigned as the end address of the current node. V. The address of each node will be concatenated by moving in the reverse order to the right (larger) direction: if the end address of the previous node ([n-1]) is smaller than the address of the current node ([n]), the address of the previous node will be assigned as the address of the current node minus the output feature map size of the previous node.

[0026] Figure 1 An exemplary flow chart of the process of allocating memory addresses to neural network nodes is presented. The proposed method manages memory allocation by classifying neural network nodes into two categories: ordinary nodes and spliced ​​nodes. This classification process is marked as step S100, which involves analyzing deep neural network nodes and distinguishing them according to their functions and structures. By classifying nodes into two categories, ordinary nodes and spliced ​​nodes, the proposed method is able to implement a more efficient and customized memory management strategy. Next, the ordinary nodes are arranged in order from large to small according to the size of their feature maps (step S102). The proposed method ensures that in the subsequent memory allocation process, the nodes with the largest feature maps are given priority.

[0027] After the neural network nodes are divided, the common nodes are arranged in descending order according to the size of their feature maps (step S102). Subsequently, memory addresses are assigned to these common nodes through the best match search technique (step S104). The best match search strategy ensures that the available memory space can be effectively utilized by finding the smallest memory block that can meet the needs of each common node. Finally, a loose match search method is used to assign memory addresses to the spliced ​​nodes (step S106).

[0028] The memory planning algorithm allocates memory efficiently for each node's intermediate feature graph by prioritizing memory allocation in SRAM and then in DRAM. The algorithm strictly adheres to the predetermined size and life cycle limits of SRAM and DRAM, allowing SRAM to avoid processing nodes with too long a life cycle.

[0029] The best match search algorithm tries to identify existing nodes that have been allocated addresses and whose lifetime overlaps with the current node. Then, it locates the memory spaces created by these nodes and selects the smallest space that can accommodate the current node. The proposed algorithm assigns the starting address of this space to the current node and adds the node to the list of allocated nodes, sorting the list in ascending order of address.

[0030] The relaxed matching search algorithm divides the concatenated tuples into two lists: a list of concatenated tuples of unique nodes (subcategory 1), where any concatenated tuple in the list has no shared / common nodes with other tuples and can be allocated in SRAM or DRAM, and a list of concatenated tuples with common nodes (subtype 2). For each tuple in the list of concatenated tuples of unique nodes, all nodes in the same tuple can be combined into a supernode. The start node of the supernode is the minimum start node of all individual nodes in the tuple, the end node is the maximum end node among them, and its size is the cumulative feature map size of all nodes in the tuple. The supernode is then processed using the best matching search algorithm to determine its address in memory. The address of the first node in the tuple is set to the address of the supernode. For subsequent nodes ([n+1]), their addresses are determined by the end address of the current node ([n]), that is, the address of the current node plus its output feature map size. All nodes in the same tuple are then inserted into the allocated node list, and the list is sorted in ascending order of address.

[0031] For each tuple in the list of concatenated tuples with common nodes, all nodes in the tuple are initially assigned an address that is immediately after the largest end address of all previously assigned overlapping nodes. All nodes are then inserted into the list of assigned nodes and sorted in ascending order of address. The address of each node is then adjusted to maintain its order in the concatenation operator: if the address of the next node is less than the end address of the current node, the address of the next node is updated to the end address of the current node.

[0032] The address of each node is reorganized by moving the nodes to the right in reverse order. Specifically, if the end address of the previous node (n-1) is less than the start address of the current node (n), the end address of the previous node is updated to the start address of the current node minus the output feature map size of the previous node.

[0033] Figure 2 Another method for enhancing memory efficiency when executing deep neural networks on resource-limited edge devices is presented. The proposed method facilitates more efficient memory management by organizing the nodes at the input deep neural network (DNN) model graph 200 into different categories to simplify neural computation. After receiving the DNN model graph, at 202, the system identifies the nodes and classifies them into two main categories: a list of ordinary nodes at 206 and a list of concatenated node tuples at 204. At 208, the ordinary nodes are sorted in descending order according to the size of the feature graphs they generate in order to implement a priority allocation strategy. Thereafter, at 210, a best match search algorithm is used to allocate appropriate memory addresses for these nodes within static random access memory (SRAM) and dynamic random access memory (DRAM) cells, and the allocation process follows the predefined size and life cycle limits of the nodes. The SRAM size limit is set to determine the maximum amount of SRAM that can be used for memory allocation, thereby ensuring that the SRAM is neither overused nor underutilized. By parsing the SRAM capacity, the proposed method can effectively allocate memory addresses between ordinary nodes within the SRAM limit, thereby optimizing fast access time and reducing latency.

[0034] The concatenated nodes from 204 are processed at 222 using a relaxed matching search algorithm, which allows for greater flexibility in allocating memory addresses due to the inherent merged output characteristics. By allocating memory addresses in this customized manner, the proposed method effectively utilizes the limited memory resources of edge devices, improving their ability to run complex DNNs while maintaining system reliability and processing speed.

[0035] After the DNN nodes are initially classified into common nodes and concatenated nodes according to their different characteristics, the "get common node list" operation at 206 involves extracting all nodes classified as common nodes from the input DNN model graph 200. The extraction here is part of the preprocessing stage and is the main step before further processing is performed on the common nodes, which can lay the foundation for the subsequent memory allocation procedure so as to efficiently execute the DNN on the edge device.

[0036] The proposed method ensures that larger nodes are prioritized for memory placement, thereby achieving more efficient memory distribution by arranging regular nodes in descending order according to the size of the feature map. This arrangement seeks to minimize the space waste that may occur when trying to fit larger nodes into the gaps between smaller nodes. Continuing this strategic direction, static random access memory (SRAM) and dynamic random access memory (DRAM) can be effectively utilized, further improving the overall efficiency of DNN execution in the constrained environment of edge devices.

[0037] Through the proposed classification and sorting methods, a certain implementation method can adopt different memory allocation strategies suitable for various node characteristics, such as the best matching search strategy for ordinary nodes and the loose matching strategy for splicing nodes, thereby optimizing the use and management of precious and limited memory resources in edge computing applications.

[0038] Regarding ordinary nodes, Figure 2 The method uses a best fit search algorithm that checks the available memory space and matches these ordinary nodes to the space in a manner similar to the best match, thereby minimizing the waste of memory resources as much as possible. At 212, the best match search algorithm assigns memory addresses to ordinary nodes in DRAM, ensuring that each node obtains an address that best matches its size requirements, thereby helping to reduce memory fragmentation and ensuring more uniform use of DRAM. This careful alignment of nodes to memory space not only improves memory utilization, but also enhances the performance of DNNs running on edge devices, which is a critical feature given that such devices are typically resource-limited. The results of the run are saved in a node repository of allocated memory resources at 214.

[0039] Due to the nature of the splicing nodes themselves, their memory management method is different from that of ordinary nodes. The relaxed Fit Search algorithm is designed to provide flexibility in the memory allocation process of these splicing nodes. The algorithm identifies suitable memory spaces in SRAM that can meet the variable size requirements of splicing nodes without imposing strict constraints unique to the best matching strategy. Applying this algorithm ensures that memory fragmentation is minimized and the utilization of available SRAM is optimized, both of which are very important in resource-constrained edge devices.

[0040] The spliced ​​nodes are processed by a relaxed matching search algorithm to allocate addresses in the SRAM at 222, thereby achieving greater flexibility within the constraints of memory management and node requirements. At 224, the remaining spliced ​​tuple list is used to remap the allocated addresses back to the nodes in the input DNN model graph at 200, thus ensuring that the memory allocation can be seamlessly integrated in the neural network structure while improving the efficiency and performance of the edge device.

[0041] The node's lifecycle limit is evaluated to determine the duration that the node can remain relevant during processing of the DNN. This parameter is used to provide nodes with shorter lifecycles that can benefit from faster SRAM memory, while nodes with longer lifecycles may be more suitable for allocation to dynamic random access memory (DRAM) (step 226). By evaluating the lifecycle of each node and juxtaposing it with the SRAM size limit, the method of the present invention ensures that memory allocation is both strategic and economical, thereby maximizing the potential of limited memory resources available for DNN computations on edge devices.

[0042] After completing the classification and sorting of the nodes in the deep neural network, the edge device integrates the newly allocated memory addresses into the network structure for seamless operation. After completing the allocation of ordinary nodes and spliced ​​nodes, each node and its assigned address will be compiled into the allocated node list 214. The proposed method continues at 220 to assign the address assigned to each node back to the network model graph. The address assigned to each node is reassigned back to the network model to update the internal data flow of the neural network. Therefore, each node in the input deep neural network model graph 200, whether it is a normal node processed by steps 210 to 213 or a spliced ​​node processed by steps 222 to 226, is now linked to a specific memory address so that the neural network operation can be performed accordingly. This processing procedure can allocate optimized memory allocation slots to corresponding nodes in the neural network, enabling the device to effectively manage cognitive tasks without additional memory access delays or computing bottlenecks, thereby maximizing the use of the limited computing resources of the edge device.

[0043] The proposed method provides an innovative solution for managing dynamic memory allocation of deep neural network (DNN) nodes in edge devices. The proposed method extends memory allocation by distinguishing between "normal" nodes and "spliced" nodes in the DNN architecture. Normal nodes, identified by the system as nodes without splicing operations, are sorted in descending order by the size of their feature maps. Spliced ​​nodes, on the other hand, are managed separately due to their different memory characteristics, where spliced ​​nodes are normal nodes connected to a splicing operator that merges feature maps from these nodes. After sorting, normal nodes undergo a meticulous memory allocation process to optimize memory usage. This process is achieved by employing a best match search algorithm that finds the most suitable memory space in static random access memory (SRAM) and dynamic random access memory (DRAM) based on preset size and life cycle constraints. SRAM is faster but usually smaller in capacity and is considered a premium resource, which is preferentially allocated to nodes that require more immediate and fast access, especially those with smaller feature maps and shorter life cycles.

[0044] Figure 3 More details of the operation of the best match search process are presented. The proposed method is further refined and begins by retrieving the list of allocated nodes 310. For each allocated node, its start and end nodes and size are obtained at 312, allowing the maximum start node and minimum end node to be calculated at 314. The system checks if there is overlap in the node life cycle at 316 and adds any overlapping nodes to the current node's overlapping node list 318. This check is repeated until there are no nodes to be verified at 320.

[0045] At the same time, the unallocated nodes are processed. At 330, a list of unallocated nodes is obtained, and detailed information of each unallocated node, such as a starting node, an ending node, and a size, is obtained at 332. At 334, it is determined whether the life cycle of the node exceeds the life cycle limit. At 342, it is checked whether there are overlapping nodes at the current node, and at 344, the memory space divided by these overlapping nodes is identified.

[0046] At 346, the smallest suitable space that can accommodate the current node is selected. Once a suitable space is found at 348, the address of the current node can be allocated accordingly at 350 and added to the list of allocated nodes at 354. Finally, the list of allocated nodes is sorted in ascending order according to the allocated address at 356. This process continues in a loop until all nodes are allocated at 358.

[0047] After obtaining the list of allocated nodes at 310, the next step in the method involves obtaining the start node, end node, and size of each allocated node at 312. This operation is very important for accurately determining the location and size attributes of each node in the allocated memory space. The collected information is then used to identify the maximum start node and the minimum end node to provide support for subsequent steps in the memory management process.

[0048] At 316, after determining whether the lifecycles of the nodes overlap ("Check: Maximum Start Node ≤ Minimum End Node"), the flow proceeds to step 318. Here, if the lifecycles do overlap, the allocated node is added to the list of overlapping nodes associated with the current node. This step helps manage and track nodes that share overlapping lifecycles, ensuring efficient use of memory resources during execution of neural networks on edge devices.

[0049] At 332, the handler obtains the start node, end node, and size of each unallocated node (called the current node). This step determines whether the node can be accommodated within the specified life cycle limit and is critical for effectively managing memory resources. At 334, in the handler, the evaluation condition is "Is (end node - start node) greater than the life cycle limit?" This decision point evaluates whether the difference between the end node and the start node of each unallocated node exceeds the predetermined life cycle limit (LifeCycleLimit).

[0050] Starting from 320, the processing program determines the overlapping node list of the current node at 342 (this list may be empty). In this step, after determining whether the difference between the end node and the start node of the unallocated node exceeds the life cycle limit, the processing program retrieves a list of nodes that overlap with the current unallocated node under consideration in memory space. This overlapping node list helps to identify the possible memory space (gap) divided by these nodes, so that in the subsequent steps, the address space can be allocated to the unallocated node.

[0051] The handler first obtains a list of overlapping nodes for the current node, which may be empty. The subsequent steps at 344 involve identifying the memory spaces or gaps delimited by these overlapping nodes. These gaps represent potential memory areas that can be used for allocation. By identifying these gaps, the memory allocation process can be optimized because the process can ensure that the available memory is used efficiently and nodes are correctly allocated to the appropriate memory locations.

[0052] Next, at 348, the handler checks if there is free space, which is a decision point in the memory allocation process when the edge device manages neural network nodes. This step involves verifying if there is a memory space of appropriate size available to accommodate the currently unallocated node. If such a space is found, the handler assigns the address of that space to the current node. However, if no suitable space exists, alternative steps must be taken to ensure efficient memory management.

[0053] At 350, the address of the current node is set to the address of the available memory space. This is part of the memory allocation process, which searches to find the smallest gap large enough to accommodate the current node. If there is free space, the address of the current node is assigned to this identified space. This ensures that memory is allocated efficiently, fragmentation is avoided, and the available memory resources are maximized.

[0054] Next, at 354, the handler pushes the current node into the list of allocated nodes. This operation adds the current node that has successfully allocated a memory address to the list of allocated nodes. This step ensures that the node is identified as allocated and properly tracked throughout the memory management process. In this way, the system can maintain an ordered record of all allocated nodes, which can better support subsequent memory operations and optimizations.

[0055] Figure 4 An exemplary pseudocode for performing static memory planning is presented. The code collects the convolution (Conv) node information and initializes the node's address and SRAM allocation flag in lines 5-12. Then, the code combines the input node IDs of each splicing node into a tuple in lines 13-17 and sorts them in descending order of the output feature map size (fmoSize). A greedy algorithm (i.e., best match search algorithm) is applied to allocate SRAM memory for the Conv node, and a relaxed greedy algorithm (i.e., relaxed matching search) is used to allocate SRAM memory for the splicing node in lines 21-22. Next, for DRAM storage, a greedy algorithm is applied to allocate DRAM memory for the Conv node, and a relaxed greedy algorithm is used to allocate DRAM memory for the splicing node in lines 24-25. The calculated address is assigned back to each node in the graph.

[0056] Figure 5An exemplary pseudocode for performing a best match search is presented. Lines 2-4 are registers for initializing the last address, the best address, and the minimum space size (lines 2-4). The code imposes a lifecycle limit MAX_LIFE (i.e., LifeCycleLimit) on the node. If the node has been allocated to SRAM, or the fmo size of the node is larger than the RAM (SRAM / DRAM) size, or it has a spliced ​​output, the subsequent node processing process is skipped (lines 5-11). For the remaining nodes k, lines 17-19 of the code can determine whether node k overlaps with node r in the lifecycle, and if so, lines 20-25 of the code calculate the size of the available space (spaceSize). If this spaceSize is not less than fmoSize[k] and not greater than the previously detected minimum space size (minSpaceSize), the minimum space size (minSpaceSize) is updated, and the best address is assigned as the starting address (spaceAddr) of the space. Next, if the sum of the address of node r and its output feature map size (the address after its end address) is greater than the starting address of the space, the starting address of the space is updated accordingly. If the best address has not been allocated before, it is allocated as the latest starting address of the available space. The code checks whether the best address plus fmoSize_k (or fmoSize[k]) is still within the size limit, and then updates the maximum total memory size (lines 26-30). At line 32, the allocated nodes are sorted in ascending order of node address (nodes with smaller addresses are listed first).

[0057] At 356, the processing program can sort the list of allocated nodes in ascending order of the allocated addresses, and then check whether there are still unchecked unallocated nodes at 358. If so, the allocation process is repeated for the remaining unallocated nodes. This ensures that all nodes are efficiently allocated memory addresses, thereby optimizing the storage of intermediate feature maps of the deep neural network model.

[0058] For the case where a node exceeds the specified lifetime limit, the algorithm checks the overlapping node list to find a memory space that can accommodate the long-lived node. After the smallest suitable space is identified, the address of the current node is set to the address of the space, and then the node is added to the allocated list and sorted. Unchecked nodes continue through this cycle until all nodes are allocated.

[0059] Fig. 6A An exemplary pseudo code for performing a loose match search is shown, and Figure 6BThe corresponding flowchart is shown. The processing program can divide the common node tuple list into 2 groups: unique node list and common node list (lines 3-13). For each tuple (qTuple) in uniqueNodeTupleList (unique node list), the code can calculate (a) the minimum starting node, (b) the maximum ending node, and (c) the total fmoSize. All nodes in each qTuple form a supernode, and (a), (b) and (c) are node information (lines 19-25). The code applies the best fit search (bestFitSearch) algorithm to calculate the best address of the supernode (lines 27-43), and calculates the addresses of the nodes in qTuple (lines 44-48), and sorts them in ascending order by address (line 49). For each tuple (nTuple) in commonNodeTupleList (common node list), the code obtains a node (u) that has not been assigned an address before (lines 51-55). For each node in allocateNodeList, the code skips the nodes in nTuple (lines 60-61). For the remaining nodes, the code does the following: Calculate the address after each overlapping node. Get the maximum address of all overlapping nodes and assign it to node (u) (lines 62-73), then sort them in ascending order of address (line 74); adjust the address of each node to align with the original splicing order (lines 75-83); and, adjust the address of each node to glue them together (stitch) together (lines 84-91). Fig. 6A The code is presented in the form of a flowchart. Figure 6B .

[0060] The remaining address allocations can be performed by assigning back to the model graph, which maps each node to its designated memory address. The adopted approach ensures optimal memory usage, balancing the allocation choices between SRAM and DRAM to improve performance, reduce power consumption, and meet the strict memory constraints of edge devices. This implementation provides a powerful framework for managing memory in edge-based neural processing units (NPUs), thereby facilitating efficient storage and retrieval of feature maps in DNN applications.

[0061] Embodiments also include applying specific size limits on memory allocations in SRAM and DRAM to ensure that memory usage remains within the provided limits. These size limits help manage memory allocations efficiently, minimizing over-allocations so that other operations do not run out of memory. The proposed method utilizes best-match and relaxed-match search algorithms to achieve optimal allocation while taking into account current and future memory requirements.

[0062] By adopting the described approach, embodiments significantly improve the management of memory resources on edge devices, enhancing the ability to run multiple artificial intelligence (AI) models simultaneously while keeping memory consumption and power consumption to a minimum. This solution is particularly beneficial for devices with strict memory and power budgets, such as edge computing devices and neural processing units (NPUs).

[0063] In an exemplary embodiment, the proposed method starts with an input DNN model graph, and the DNN model graph undergoes node splitting and information extraction to classify nodes into ordinary nodes and splicing tuples. The splicing tuple is a list of nodes that need to be spliced ​​together, while the ordinary nodes are sorted in descending order of feature graph size. After extraction, the system retrieves the list of splicing tuples and the list of ordinary nodes. After the address is assigned, the current node is pushed into the list of allocated nodes. Then, the allocated nodes are sorted in ascending order of address. The system can re-evaluate to determine whether any unallocated nodes are not selected and repeat as needed. The next step is to select the minimum space or gap that is available and large enough to accommodate the current node. The algorithm used can evaluate the size requirement of the current node and compare it with the size of the identified gap. By selecting the smallest and large enough space, the algorithm used aims to optimize memory utilization and minimize fragmentation. Finally, the algorithm used assigns the starting address of the selected space to the current node. This allocation mechanism ensures that memory addresses are allocated efficiently and that available memory is fully utilized to accommodate the size and life cycle of the current node without wasting resources. Therefore, the address of the current node becomes part of the overall allocated node list, which is ready to be reintegrated into the computational graph of the deep neural network. This insertion process can be done in an orderly manner, ensuring that the addresses in the allocated node list are arranged in a predetermined order, and usually in ascending order of addresses, to facilitate further operations and memory management tasks.

[0064] In an efficient memory reuse and allocation method for a deep neural network (DNN) model graph on an edge device, a processing procedure starts with an input DNN model graph. The nodes in the graph are split and information about the start node, end node and size of each node is extracted. Two main lists can be created from these nodes: a concatenated tuple list containing nodes that need to be concatenated together and a normal node list containing the remaining nodes. Subsequently, the nodes in each list can be sorted in descending order of their feature map size. Each allocated node is then reinserted into the allocated node list, which can be sorted in ascending order according to the newly allocated memory address. For nodes in the concatenated tuple list that share common nodes with other tuples, memory addresses are specifically allocated in DRAM. The memory addresses of these nodes are allocated to be immediately after the maximum end address of all previously allocated overlapping nodes and are recalibrated to prevent them from overlapping with the allocated nodes. This allocation procedure is performed to preserve the logical order of operations in the DNN to avoid differences during graph reconstruction. After allocation, it is determined whether there are other unchecked allocated nodes. If any node is not processed, the allocation algorithm will iterate until all nodes are properly processed. The goal is still to optimize memory utilization while maintaining computational efficiency and ensuring that the node lifecycle constraints within the SRAM and DRAM limits are respected. Re-inserting each node with an allocated address into the graph ensures that the allocated memory is established in a sequential and systematic manner to run the DNN model. The detailed process includes first obtaining a list of unallocated nodes and extracting details such as their start node, end node, and size. The overlap of the lifecycles between the nodes can be evaluated again, and if there is an overlap and the lifecycle is within an acceptable range, the node is marked for SRAM allocation. Nodes with lifecycles exceeding the specified limits will be considered for DRAM allocation. The allocation of each node involves determining the space in the available memory, finding the gap formed between previously allocated and overlapping nodes, selecting the minimum suitable space for the current node, and allocating the address corresponding to this space. Through these methods and steps, an efficient memory reuse method can achieve an optimal allocation pattern of memory addresses in SRAM and DRAM, thereby reducing memory consumption and power consumption on neural processing units (NPUs) on edge devices. This innovative approach can significantly save memory space and improve the energy efficiency of DNN models deployed on these devices.

[0065] The loose matching search algorithm divides the spliced ​​tuples into two subtypes, the spliced ​​tuple list of unique nodes (subtype 1): which is a spliced ​​tuple that has no shared / common nodes with other spliced ​​tuples; and, the spliced ​​tuple list of common nodes (subtype 2): which is a spliced ​​tuple that has shared / common nodes with other spliced ​​tuples. The spliced ​​tuple list of unique nodes can be allocated in SRAM or DRAM. The spliced ​​tuple list of common nodes can only be allocated in DRAM. This implementation solves the limitation of limited memory size in edge devices and achieves efficient memory reuse, thereby optimizing the memory allocation process of deep neural network models running on neural processing units (NPUs).

[0066] In addition, the relaxed matching search algorithm divides the spliced ​​tuples into two subtypes: unique nodes that have no shared / common nodes with other tuples, and common nodes that have shared common nodes with other tuples. Unique node tuples can be allocated in SRAM or DRAM, while common node tuples can only be allocated in DRAM. For unique node tuples, all nodes in the same tuple will be combined into a supernode, whose start node is the minimum start node of all nodes in the tuple, and whose end node is the maximum end node of all nodes in the tuple. The total size of the feature graphs of all nodes in the tuple defines the size of the supernode. The supernode goes through the best matching search algorithm to obtain the address in memory. After the address is assigned, the address of the first node in the tuple will be set to the address of the supernode, and the subsequent node addresses will be assigned according to the end address of the current node plus its output feature graph size. Then, all nodes in the tuple are inserted into the assigned node list in ascending order of address.

[0067] For each common node tuple, the proposed method assigns the address immediately after the maximum end address of the previously assigned overlapping node, excluding the previously assigned common nodes. Finally, all nodes are inserted into the assigned node list and sorted in ascending order of address. This comprehensive approach optimizes the memory usage of edge devices with strict memory constraints, ensuring that the feature graphs of neural network models can be stored efficiently and in order.

[0068] In one embodiment, the method includes a memory planning algorithm that divides nodes into two categories: one is ordinary nodes, which are sorted in descending order of their feature graph sizes; the other is nodes that need to be spliced, where a group of nodes spliced ​​together is called a spliced ​​tuple.

[0069] The memory planning algorithm uses the best matching search algorithm to allocate memory addresses for common nodes, and uses the relaxed matching search algorithm to allocate memory addresses for splicing nodes.

[0070] The memory planning algorithm allocates memory for the intermediate feature graph of each node in SRAM and DRAM, first allocating memory in SRAM and then allocating memory in DRAM, and based on the provided size constraints, i.e., the SRAM size constraint and the DRAM size constraint, and based on the provided life cycle constraints of the nodes to be allocated in SRAM and the life cycle constraints of the nodes to be allocated in DRAM, so that the SRAM can avoid processing nodes with too long life cycle.

[0071] The best match search algorithm aims to efficiently allocate memory addresses to nodes. First, a list of allocated nodes that overlap with the lifetime of the current node is obtained. Then, the memory space divided by these allocated and overlapping nodes is identified. The algorithm continues to search for the smallest space that can accommodate the current node and assigns the starting address of this space to the current node. Finally, the current node is inserted into the list of allocated nodes and sorted in ascending order of address.

[0072] On the other hand, the relaxed matching search algorithm divides the splicing tuples into two subtypes: unique node tuples and common node tuples. Unique node tuples do not share nodes with other splicing tuples and can be allocated in SRAM or DRAM. Common node tuples that share nodes with other splicing tuples can only be allocated in DRAM.

[0073] For a unique node tuple, the algorithm combines all nodes in the same tuple into a supernode. The supernode's start node is the minimum start node of all nodes in the tuple, its end node is the maximum end node of all nodes in the tuple, and its size is the total size of the feature graph of all nodes. The supernode then goes through the best match search algorithm to obtain the memory address. The address of the first node in the tuple becomes the address of the supernode, and the addresses of subsequent nodes are calculated based on the end address of the previous node. All nodes in the tuple are then inserted into the allocated node list and sorted by address.

[0074] For a common node tuple, the algorithm first assigns addresses to all nodes in the tuple that are immediately after the maximum end address of the allocated overlapping nodes, but does not include the allocated common nodes. After inserting all nodes into the allocated node list and sorting them by address, the algorithm adjusts the node addresses to maintain their order in the splicing operator. Finally, the addresses are spliced ​​by moving the nodes to the right in reverse order, ensuring correct alignment and reasonable allocation of memory space.

[0075] In the process of efficiently allocating memory addresses in the DNN model graph, the addresses of the nodes are adjusted according to specific rules to optimize memory usage. If the address of the next node in the sequence is found to be less than the end address of the current node, the system adjusts by assigning the address of the next node to the end address of the current node. This adjustment helps maintain the ascending order of addresses, ensuring that each subsequent node does not overlap the memory space allocated by the previous node. This adjustment prevents conflicts and ensures that memory allocation is efficient and orderly. As part of a broader memory planning algorithm, this step plays a key role in maintaining the integrity and continuity of the memory address ranges assigned to each node in the model graph. Specifically, this also involves inserting all nodes into the list of allocated nodes and maintaining them in ascending order of addresses. With the above arrangement, the algorithm adopted can not only maximize the use of available memory, but also ensure that the nodes are arranged in a logical order. This address adjustment process helps to reduce memory fragmentation overall and allows more nodes to be accommodated in a limited memory space, thereby optimizing the performance and efficiency of the neural network processing unit. Therefore, this approach is particularly helpful for edge devices and neural processing units (NPUs) where memory resources are strictly limited and efficient reuse of memory addresses is extremely important.

[0076] In another embodiment, a method for efficiently utilizing memory space for storing intermediate feature graphs of a deep neural network model is provided. The method includes a memory planning algorithm, which divides nodes into two categories: one is ordinary nodes, which are sorted in descending order of their feature graph sizes; the other is nodes that need to be spliced, wherein a group of nodes that need to be spliced ​​together is called a splicing tuple. The memory planning algorithm uses a best match search algorithm to allocate memory addresses for ordinary nodes, and uses a loose match search algorithm to allocate memory addresses for splicing nodes. The memory planning algorithm allocates memory for the intermediate feature graph of each node in SRAM and DRAM, and first allocates memory in SRAM, and then allocates memory in DRAM, and must be based on the size limit provided, that is, the size limit of SRAM and the size limit of DRAM, and based on the life cycle limit of the node to be allocated in SRAM and the life cycle limit of the node to be allocated in DRAM, so that SRAM can avoid processing nodes with too long a life cycle. The best matching search algorithm attempts to obtain a list of nodes that have been allocated addresses and overlap with the current node in their life cycle from the allocated node list; find a list of memory spaces divided by the allocated and overlapping node lists; find the minimum space that is large enough to accommodate the current node; assign the starting address of the space as the address of the current node; insert the current node into the allocated node list, and sort the allocated node list in ascending order of address. The loose matching search algorithm divides the spliced ​​tuples into two subtypes, the spliced ​​tuple list of unique nodes (subtype 1): it is a spliced ​​tuple that has no shared / common nodes with other spliced ​​tuples, and the spliced ​​tuple list of common nodes (subtype 2): it is a spliced ​​tuple that has shared / common nodes with other spliced ​​tuples. The spliced ​​tuple list of unique nodes can be allocated in SRAM or DRAM, while the spliced ​​tuple list of common nodes can only be allocated in DRAM. For each tuple in the concatenated tuple list of unique nodes, all nodes in the same tuple will be combined into a supernode, where its start node is the minimum start node of all nodes in the tuple, its end node is the maximum end node of all nodes in the tuple, and its size is the total feature map size of all nodes in the tuple. The supernode will get an address in memory through the best match search algorithm; the address of the first node in the tuple is the address of the supernode; the address of the next node ([n+1]) is the end address of the current node ([n]): that is, the address of the current node plus the output feature map size of the current node; all nodes in the same tuple are inserted into the allocated node list and sorted in ascending order by address.For each tuple in the concatenated tuple list of common nodes, all nodes in the same tuple will first be assigned the address immediately following the largest end address among all previously assigned overlapping nodes; all previously assigned overlapping nodes do not include previously assigned common nodes; all nodes are inserted into the list of assigned nodes and sorted in ascending order by address; the address of each node will be adjusted to preserve the same order in the concatenation operator: if the address of the next node is less than the end address of the current node, the address of the next node is assigned as the end address of the current node. The addresses of each node will be concatenated together by moving the nodes to the right (larger side) in reverse order: if the end address of the previous node ([n-1]) is less than the address of the current node ([n]), the address of the previous node is assigned as the address of the current node minus the output feature map size of the previous node. This dual algorithm approach of classifying nodes and applying specific memory allocation strategies ensures optimized use of available memory to improve the processing efficiency of the NPU and may significantly reduce the operating power consumption of edge devices. Given that edge devices have to continually operate under limited memory capacity and high requirements for processing power, this approach provides a viable solution to improve the efficiency and performance of DNN models in resource-constrained environments.

[0077] The processing procedure starts with the division of the neural network nodes. These nodes are divided into two categories: normal nodes and concatenated nodes. Normal nodes are nodes that are processed independently, while concatenated nodes are groups of nodes that need to be processed together because of operational dependencies, such as concatenation of feature maps. This classification is important because it determines the subsequent steps of memory allocation. After the nodes are classified, the next step is to sort the normal nodes. The sorting algorithm arranges these nodes in descending order of feature map size. This is a critical step because during memory allocation, larger nodes need to be considered first to optimize the use of available memory space and also minimize the occurrence of fragmentation. In order to assign memory addresses to the sorted normal nodes, the device can use a best match search algorithm. The algorithm used searches for the smallest available memory block that can accommodate the current node. In this way, it can ensure efficient use of memory and avoid unnecessary waste of memory space that may occur by allocating larger blocks when smaller blocks are sufficient.

[0078] By applying a relaxed matching search algorithm on a model with stitched nodes, memory stitching of all stitched nodes is achieved. In addition, the method described in this paper can reduce power consumption by 46% and reduce DRAM storage requirements by 78%.

[0079] FIG7 is an exemplary chart showing memory savings for various machine learning models such as openpose, yolov3, handLandmark42D, srnet, resnet34, resnet50, and faceID. For each model, the table provides data on the reclaimed and non-reclaimed storage sizes. The non-reclaimed storage size refers to the total amount of memory (in bytes) required to store the feature maps of each node in the model without applying memory reuse techniques. Essentially, this is the sum of the memory used by all feature maps in the entire model node / layer without any optimization of memory efficiency. On the other hand, the reclaimed storage size is the amount of memory required to store the feature maps after applying the proposed memory reuse algorithm. The algorithm optimizes memory usage by reusing memory space that is no longer required by previous layers, thereby reducing the overall memory footprint of the model. The compression ratio is calculated by dividing the reclaimed storage size by the non-reclaimed storage size. This ratio reflects the efficiency of memory space compression. Among the listed models, handLandmark42D has the highest compression efficiency of 4.71%, while resnet50 has the lowest compression ratio of 20.04%. The storage saving rate is an additional metric to evaluate the efficiency of memory reuse. It can be determined by subtracting the compression ratio from 100%, indicating the percentage of storage saved. This metric ranges from 77.81% for openpose to 95.29% for handLandmark42D. The theoretical lower limit for each model represents the minimum storage requirement that can be achieved, and no algorithm can exceed this limit. The algorithm reaches its corresponding theoretical lower limit in five of the seven models, except for srnet and resnet50. In addition, the SRAM size is configured to be approximately one-tenth of the DRAM recycling size, ranging from 24K for faceID to 2048K for srnet. With this configuration, significant power consumption reduction can be achieved, with openpose achieving the largest power consumption reduction ratio of 45.5% and faceID achieving the lowest power consumption reduction ratio of 13.48%.

[0080] Although multiple alternative embodiments of the present invention have been shown, it will be appreciated that those skilled in the art may make certain modifications without departing from the basic scope of the present invention as described above and below. In addition, the above embodiments are only intended to illustrate the principles of the present invention and are not intended to limit the scope of the present invention to the disclosed elements.

Claims

1. A method for storing one or more neural network model feature graphs, comprising: Divide the neural network nodes into common nodes and splicing nodes; Sorting the common nodes in descending order of feature graph size; Using best match search, assigning a memory address to the common node; and Using a loose matching search, a memory address is assigned to the splicing node.

2. The method according to claim 1 also includes allocating memory among static random access memory (SRAM) and dynamic random access memory (DRAM), wherein the memory is first allocated in the SRAM and then allocated in the DRAM.

3. The method of claim 2, further comprising applying size restrictions to the memory allocations in the SRAM and the DRAM. 4 . The method of claim 2 , further comprising applying a lifecycle restriction to nodes to be allocated in the SRAM and the DRAM. 5 . The method according to claim 4 , comprising avoiding allocating nodes having a preset life cycle into the SRAM.

6. The method of claim 1, wherein the best match search comprises: Identifying allocated nodes with overlapping lifecycles; Determining a memory space divided by the allocated nodes; Select a predetermined memory space to save the current node; as well as Assigning the starting address of the selected memory space to the current node.

7. The method of claim 6, further comprising inserting the current node into an allocated node list and sorting the allocated node list in ascending order of address.

8. The method of claim 1, wherein the loose match search comprises: Divide the splicing nodes into unique node tuples and common node tuples; Allocate unique node tuples in SRAM or DRAM; as well as Only common node tuples are allocated in DRAM.

9. The method of claim 8, further comprising, for each unique node tuple, combining nodes into a supernode.

10. The method of claim 9, wherein the supernode comprises: A starting node, which is the minimum starting node of all nodes in the tuple; an end node, which is the maximum end node of all nodes in the tuple; as well as Size, which is the total size of the feature graphs of all nodes in the tuple.

11. The method of claim 9, further comprising applying the best match search to allocate memory for the supernode.

12. The method according to claim 11, further comprising: assigning a supernode address to a starting node in the tuple; Allocate the address of subsequent nodes according to the end address of the previous node; as well as Insert the nodes in the tuple into the list of allocated nodes, sorted in ascending order by address.

13. The method according to claim 8, further comprising: For a common node tuple, assigning its address to be immediately after the predetermined end address of the previously assigned overlapping node and excluding the previously assigned common node; as well as Insert the node into the list of allocated nodes and sort the nodes in ascending order by address.

14. The method according to claim 13, further comprising: adjusting node addresses in the common node tuple to preserve order in a concatenation operator, wherein if an address of a next node is less than the end address of a current node, the next node address is assigned as the end address of the current node; and The node addresses are concatenated by moving the nodes in reverse order, and if the end address of the previous node is smaller than the address of the current node, the address of the previous node is set to the address of the current node minus the output feature map size of the previous node.

15. A method for storing a neural network model using a memory planning algorithm, comprising: Classifying neural network nodes into first-category nodes and second-category nodes to be spliced; Sort the first type of nodes in descending order of feature graph size; Combining the second type of nodes to be spliced ​​together into a splicing tuple; Allocate memory for the intermediate feature map of each node; Allocating memory addresses for the first type of nodes using a best match search; and A memory address is assigned to the splice node using a relaxed matching search.

16. The method according to claim 15, wherein the memory planning comprises allocating memory for the intermediate feature graph of each node in the following manner: Allocate memory between static random access memory (SRAM) and dynamic random access memory (DRAM); Prioritize memory allocation in the SRAM before allocating memory in the DRAM; Comply with preset size limits of the SRAM and the DRAM; respecting preset lifecycle limits of nodes to be allocated in said SRAM and said DRAM; as well as This allows the SRAM to avoid processing nodes that have an excessively long life cycle.

17. The method according to claim 15, wherein the best match search comprises: Get a list of allocated nodes whose lifecycles overlap with the current node; identifying a memory space partitioned by the allocated and overlapping node lists; Select the minimum space that can accommodate the current node; Assigning the starting address of the selected space to the current node; as well as The current node is inserted into a list of allocated nodes and the list is sorted in ascending order of address.

18. The method of claim 15, wherein the loose match search comprises: Divide the concatenated tuple into: a list of concatenated tuples of a unique node, wherein the unique node has no shared nodes with other concatenated tuples and can be allocated in the SRAM or the DRAM; as well as a list of concatenated tuples of common nodes, which have shared nodes with other concatenated tuples and can only be allocated in the DRAM; For the tuples in the concatenated tuple list of unique nodes: combining nodes in the tuple into a supernode, wherein a start node is the start node of all nodes in the tuple, an end node is the end node of all nodes in the tuple, and a supernode size is a total feature map size of all nodes in the tuple; Applying a best match search to obtain the address of the supernode; assigning addresses to nodes in the tuple according to the addresses of the supernode; as well as inserting all nodes in the tuple into a list of allocated nodes, sorted in ascending order of address; and For each tuple in the concatenated tuple list of common nodes: Allocate addresses immediately following the largest ending address of previously allocated overlapping nodes and excluding previously allocated common nodes; Insert all nodes into the allocated node list and sort them in ascending order of address; Adjusting node addresses to preserve order in concatenation operators; and Concatenate the node addresses by moving the nodes in reverse order.

19. A system comprising: processor; One or more neural network models; as well as A module for storing feature maps of one or more neural network models, including code to: Divide the neural network nodes into common nodes and splicing nodes; Sorting the common nodes in descending order of feature graph size; Assigning a memory address to the common node using a best match search; and A memory address is assigned to the splice node using a relaxed matching search.

20. The system of claim 19, comprising a static random access memory (SRAM) and a dynamic random access memory (DRAM) coupled to the processor, wherein: The memory is first allocated to the SRAM and then to the DRAM. The processor imposes size restrictions on the memory allocations of the SRAM and the DRAM and imposes lifecycle restrictions on the nodes allocated to the SRAM and the DRAM.