Memory offset calculation method for memory overlapping degree based on tensor

By optimizing the memory allocation strategy through a tensor-based method for calculating memory overlap, the memory management problem of deep learning compilers on mobile and embedded devices is solved, achieving more efficient memory usage and computational efficiency.

CN120849083APending Publication Date: 2025-10-28HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410521457.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing deep learning compilers are limited by computing resources, storage space, and energy consumption on mobile and embedded devices, and existing memory management technologies have failed to effectively solve the memory usage problem, resulting in memory fragmentation and inefficiency.

Method used

A memory offset calculation method based on the memory overlap of tensors is adopted. By sorting the memory overlap of tensors and searching for memory gaps, the memory allocation strategy is optimized, memory fragmentation is reduced, and memory utilization efficiency is improved.

Benefits of technology

It effectively reduces peak memory size, lowers memory usage, improves computing efficiency and resource utilization, and optimizes memory management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849083A_ABST
    Figure CN120849083A_ABST
Patent Text Reader

Abstract

The invention provides a tensor-based memory offset calculation method for a memory overlapping degree. The tensor-based memory offset calculation method comprises the following steps: S1, sorting distribution priorities of intermediate tensors of a calculation graph according to the memory overlapping degree of the tensors; s2, searching a memory gap, and searching a corresponding intermediate tensor matched with a proper memory gap; and S3, overall memory allocation scheduling: performing complete scheduling of memory allocation on the neural network topological structure based on the offset calculation strategy of the memory overlapping degree of the tensor in the step S1 and the step S2. According to the method, the memory overlapping degree of the tensor is introduced into a memory offset calculation strategy, memory fragments can be solved to a certain extent, the size of a memory peak value is reduced, memory occupation of an inference engine is minimized, and memory gaps can be reduced to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning compiler optimization technology, and specifically relates to a method for calculating memory offset based on the degree of memory overlap of tensors. Background Art

[0002] With the widespread application of deep learning technology in various fields, including computer vision and natural language processing, and the latest advancements in computing hardware, the inference tasks of deep neural networks are no longer limited to the server side. More and more deep neural networks are allowing inference tasks to be transferred to mobile and embedded devices.

[0003] At the same time, the application of deep learning tasks on mobile and embedded devices has led to the following challenges:

[0004] 1) Computational resource limitations: Mobile and embedded devices typically have limited computing resources, such as CPU, GPU, and memory. This can lead to insufficient computing power during the inference process of deep learning models.

[0005] 2) Storage space limitations: Mobile and embedded devices also have relatively limited storage space, which limits the size and number of parameters of deep learning models.

[0006] 3) Energy Consumption Constraints: Mobile and embedded devices typically rely on batteries for power, thus having high energy consumption requirements. Deep learning models, on the other hand, often require significant computational and storage resources for inference, consuming substantial amounts of energy. Therefore, deep learning models need to minimize energy consumption while maintaining performance.

[0007] To address these challenges, deep learning compilers can optimize the computation graph, which helps reduce memory usage and accelerate computation. An optimized computation graph can better utilize hardware resources, such as GPUs and CPUs, thereby reducing memory accesses and improving computational efficiency. Secondly, deep learning compilers can leverage memory optimization techniques to reduce memory usage. For example, compilers can manage memory through techniques such as memory sharing and memory reclamation, avoiding problems such as memory leaks and memory overflows.

[0008] The goal of static memory planning optimization is to reuse memory buffers as much as possible. There are generally two methods: in-place memory sharing and standard memory sharing. In-place memory sharing uses the same memory for both input and output operations and allocates only one memory location before computation. Standard memory sharing reuses non-overlapping memory from previous operations. Static memory planning is done offline, which allows for the application of more complex planning algorithms.

[0009] Therefore, effectively managing the memory of deep neural networks is crucial.

[0010] However, MXNet uses heuristic algorithms (such as in-place operators and intermediate tensor memory sharing) for memory allocation, but does not explore how to solve the core memory problem, namely minimizing the memory footprint of the inference engine.

[0011] The TensorFlow Lite GPU inference engine uses a memory manager for its GPU buffers and explores two approaches to solving the memory allocation problem as a two-dimensional bar wrapping problem. However, the exploration is insufficient and performance limitations still exist.

[0012] In addition, commonly used technical terms in the prior art include:

[0013] Deep learning: Deep learning is a new research direction in the field of machine learning. It mainly learns features and patterns automatically by learning the inherent laws and representation levels of sample data, thereby enabling machines to have analytical and learning capabilities like humans.

[0014] The ultimate goal of deep learning is to enable machines to recognize and process data such as text, images, and sound, solve complex pattern recognition problems, and mimic human activities such as sight, hearing, and thinking. Deep learning compilers are specialized compilers whose primary task is to generate efficient code implementations of deep learning models on different hardware platforms. They are highly optimized for model specifications and hardware architecture, from model definition to code generation (specifically, they combine deep learning-related optimizations, such as operator fusion, to achieve efficient code generation).

[0015] Static memory planning: Determine the memory addresses and sizes of all intermediate calculation results at compile time, and through reasonable memory planning, make intermediate tensors reuse memory buffers as much as possible. Summary of the Invention

[0016] To address the aforementioned issues, this application aims to provide a reference and basis for measuring the degree of memory overlap between tensors, while also generally reflecting the extent of memory consumption during the lifetime of a tensor. Therefore, this method explores a calculation offset strategy based on the memory overlap of tensors, which can reduce the existence of memory gaps to some extent.

[0017] This method, in terms of its overall concept:

[0018] Introducing the memory overlap of tensors into the memory offset calculation strategy can, to some extent, solve the problem of memory fragmentation, reduce the size of memory peaks, and minimize the memory usage of the inference engine.

[0019] Specifically, this invention provides a memory offset calculation method based on the degree of memory overlap of tensors. For static memory planning, the method includes the following steps:

[0020] S1, sort the allocation priority of intermediate tensors in the computation graph based on the degree of memory overlap of tensors;

[0021] The sorting of the memory overlap of the tensors further includes:

[0022] The rule for ranking tensor memory overlap is to sort them in a non-increasing manner according to the following indicators: tensor overlap, tensor size, and starting point. That is, when evaluating the memory overlap of tensors, the three indicators are tensor overlap, tensor size, and starting point, with their importance decreasing in that order. S2, memory gap search, finds a suitable memory gap and matches the corresponding intermediate tensor; further including:

[0023] For each intermediate tensor, first check the allocated tensors whose lifetimes intersect with the current tensor to find the smallest memory gap between them, so that the current tensor fits the gap; if such a gap is found, the current tensor is allocated to this gap; otherwise, after allocating the current tensor to the bottommost tensor whose lifetime intersects with its own, the corresponding offset is assigned to the current tensor, and the tensor is in the allocated state.

[0024] S3, overall memory allocation and scheduling:

[0025] The offset calculation strategy based on the memory overlap of the tensors in steps S1 and S2 is used to perform a complete scheduling of memory allocation for the neural network topology. The complete scheduling process of memory allocation is as follows: S3.1, Resource initialization:

[0026] 1. Construct the allocated list:

[0027] Before any new memory allocation begins, an allocated list needs to be maintained, recording all tensors that have been allocated memory in the current system and their related information. This list typically includes information such as the tensor's identifier, starting address, required memory size, and lifetime. The purpose of building this list is to ensure that when making new memory allocations, it is clear which memory regions are already occupied, thereby avoiding memory conflicts and overlaps.

[0028] 2. Construct a list of overlapping activity levels:

[0029] The activity overlap list is used to record allocated tensors whose lifecycles intersect with the tensor to be allocated. By constructing this list, tensors that may cause memory conflicts with the tensor to be allocated can be quickly located, thus enabling more accurate memory allocation decisions. The method for constructing this list is usually to traverse the allocated list, compare the lifecycle of each allocated tensor with the lifecycle of the tensor to be allocated, and find the tensors that have intersections.

[0030] S3.2, Allocation and Scheduling:

[0031] 1. Calculate the memory required for the tensor to be allocated:

[0032] Based on the data type, data arrangement, and dimensions of the tensor to be allocated, the required memory size can be calculated; this size will serve as the basic basis for memory allocation.

[0033] 2. Retrieve memory gaps:

[0034] After obtaining information about the memory size required for the tensor to be allocated, the next step is to search for available memory gaps in the allocated list. This is typically done by iterating through the tensors in the active overlap list, calculating the memory gap size between them, and finding the gap that meets the requirements of the tensor to be allocated.

[0035] 3. Select appropriate memory gaps:

[0036] After retrieving multiple available memory gaps, it is necessary to select the most suitable gap according to a certain strategy. This strategy includes the best fit principle, that is, selecting a gap whose size is closest to the memory required by the tensor to be allocated and leaving a certain margin, so as to make full use of memory resources and reduce fragmentation.

[0037] 4. Allocate memory:

[0038] Once a suitable memory gap is selected, the tensor to be allocated can be allocated into that gap; this involves updating the starting address and memory usage of the tensor to be allocated, and adding the corresponding record to the allocated list.

[0039] 5. Handling insufficient memory situations:

[0040] If a suitable memory gap cannot be found in the active overlap list to meet the needs of the tensor to be allocated, measures to expand the memory space are needed to handle the insufficient memory situation.

[0041] 6. Update the assigned list and the list of overlapping activity levels:

[0042] After memory allocation is complete, the allocated list and the active overlap list need to be updated to reflect the latest memory usage; this includes adding newly allocated tensors to the allocated list and updating the active overlap list as needed.

[0043] The above steps enable resource initialization and allocation scheduling, ensuring that tensor memory allocation is both efficient and accurate. This process can be adjusted and optimized according to specific system requirements and optimization goals to achieve better performance and resource utilization.

[0044] In step S1

[0045] First, the tensor overlap is the most critical indicator for evaluating the degree of memory overlap, directly reflecting the degree of overlap between different tensors in memory, that is, the size of the memory space they share. The higher the overlap, the higher the memory utilization efficiency, but it may also increase the complexity of memory management. Therefore, tensors with higher overlap should be given priority in the sorting process in order to better optimize memory utilization.

[0046] Secondly, the tensor size is also an important indicator; it determines the amount of space a tensor occupies in memory. Generally speaking, larger tensors require more memory, so they should be given due consideration during sorting. By considering the tensor size, we can better balance memory usage and performance requirements, and avoid situations where memory is insufficient due to an excessively large tensor.

[0047] Finally, although the starting point is relatively minor, it can still affect the degree of memory overlap in some cases. The starting point determines the initial position of the tensor in memory, and different starting points may lead to different memory layouts and overlaps. Therefore, the influence of the starting point should also be considered appropriately during the sorting process to ensure the rationality of the memory layout.

[0048] In summary, when evaluating the memory overlap of tensors, three indicators should be considered comprehensively: tensor overlap, tensor size, and starting point, and ranked in order of their importance. Through reasonable ranking and memory management strategies, memory utilization efficiency can be improved, memory fragmentation can be reduced, and thus the overall performance of the model can be optimized.

[0049] In step S1, the intermediate tensor of the computation graph has the following characteristics:

[0050] Lifecycle: The lifecycle of a tensor represents the entire process during which the tensor is active in memory;

[0051] Memory scheduling state: The memory scheduling state represents the state of the tensor in memory throughout the entire inference timeline;

[0052] Tensor overlap: Tensor overlap represents the degree to which tensor memory states overlap during their active periods. Tensor overlap is a key indicator for measuring the degree to which different tensors share memory space, and it is influenced by two core factors: overlap size and the number of overlapping tensors.

[0053] First, the overlap size directly reflects the total size of other tensors that intersect with it during the tensor's lifetime; the calculation of the overlap size usually involves a precise analysis of the tensor's layout and location in memory to determine the overlapping area between them.

[0054] Secondly, the number of overlapping tensors is also an important factor affecting tensor overlap. It represents the number of active tensors that exist simultaneously in memory. When the number of overlapping tensors increases, it means that more tensors are active in the memory space, which helps to reduce memory fragmentation and waste. However, as the number of overlapping tensors increases, the complexity of memory management may also increase accordingly, requiring more refined strategies to ensure coordination and synchronization between the tensors.

[0055] Taking into account both the overlap size and the number of overlap tensors allows for a more comprehensive evaluation of tensor overlap. In practical applications, optimizing tensor layout and allocation strategies can maximize the overlap size and number of overlap tensors, thereby improving memory utilization efficiency. However, it's also crucial to consider the complexity and performance overhead of memory management to ensure that improving memory utilization efficiency doesn't introduce excessive additional burdens.

[0056] In summary, tensor overlap is a comprehensive indicator reflecting the degree of tensor memory sharing, influenced by both overlap size and the number of overlapping tensors. By appropriately optimizing these two factors, we can achieve more efficient memory usage and management.

[0057] The features of the intermediate tensor further include:

[0058] The lifetime of an intermediate tensor t can be defined as {start point, end point}, where the start point and end point are the producer operator of the intermediate tensor t and the index of the last operator that took the intermediate tensor t as its input, respectively. These indices come from a topological sorting of the neural network, which is also the execution order of the operators. It is worth noting that no two tensors with intersecting intervals can share memory.

[0059] The memory scheduling state of an intermediate tensor t can be defined as {activity, tensor size}, where activity represents whether the tensor exists in memory, and tensor size is the amount of space occupied by the data in the tensor in memory, in bytes.

[0060] In step S1, the degree of memory overlap is assumed to be as shown in the following table:

[0061] The order of memory overlap of tensors is: 2->5->4->3->1->0->6;

[0062] During the tensor memory overlap sorting process, tensors are first sorted according to their overlap to ensure that tensors with smaller overlap are processed first. As shown in the table, Tensor 2 has an overlap size of 54 and 5 overlapping tensors, while Tensor 3 has an overlap size of 30 and 3 overlapping tensors. Therefore, Tensor 2 has an overlap of 0, and Tensor 3 has an overlap of 2. Consequently, after sorting, Tensor 2 comes before Tensor 3. Then, among tensors with the same overlap, tensors with larger sizes are processed first. As shown in the table, Tensor 3 has a tensor size of 5, and Tensor 4 has a tensor size of 7. Therefore, after sorting, Tensor 4 comes before Tensor 3. Finally, if the first two criteria are the same, tensors with earlier starting points are processed first.

[0063] In step S2, during the search for memory gaps:

[0064] Memory gaps can be viewed as memory buffers that have been used before a certain point in time and are currently in a freed state.

[0065] In step S2, the complete memory gap search process is described as follows: When processing the intermediate tensor T of the memory to be allocated... a The following design steps are required:

[0066] First, it will be based on tensor T a By analyzing key information such as data type, data layout, and dimensions, the required memory size S can be accurately calculated. a This step is crucial, as it ensures an accurate understanding of memory requirements and lays a solid foundation for subsequent memory allocation.

[0067] Next, the list of allocated tensors will be carefully traversed to identify those intermediate tensors T that are related to the current tensor T to be allocated in terms of lifetime. a Tensor sets V with intersection t To manage these tensors more efficiently, they will be sorted in ascending order according to their starting addresses, resulting in a clearer and more orderly memory layout. Subsequently, the tensor set V will be used... t To retrieve the memory gap list V f By comparing the size of each memory gap, denoted as f0, f1, ..., f n It will search for the first tensor T whose size is greater than or equal to the tensor T. a The required memory and the most suitable memory gap; in this process, the best fit principle is followed, that is, min(f x -S a If a suitable memory gap V is successfully found, its value is greater than or equal to 0. x It will then be assigned to tensor T.a Make its starting address the same as V x The starting addresses are the same;

[0068] However, if a sufficiently large memory gap cannot be found to satisfy tensor T a If the requirement is met, the memory expansion function will be activated; at this time, the tensor T will be... a Assigned to the already assigned tensor set V t The last tensor T b After that, ensure that the starting address of tensor Ta immediately follows T. b Afterwards, and considering T b Required memory size;

[0069] Finally, to ensure the consistency and integrity of memory management, the intermediate tensor T of successfully allocated memory will be... a Add it to the list of assigned tensors for later management and querying;

[0070] By carefully designing and executing the above steps, we can ensure efficient memory utilization and reasonable allocation, thus providing a basis for tensor T. a This provides a strong guarantee for the smooth operation of the system.

[0071] In step S3, the offset calculation strategy based on the memory overlap of tensors includes: pre-allocating a large block of memory, and the intermediate tensor being a data buffer partitioned out by the offset within the memory block. This method is called memory offset calculation, and its goal is to minimize the size of the allocated memory block.

[0072] The offset calculation problem can be viewed as a special case of the two-dimensional strip filling problem. This problem can be abstracted and simplified into a filling problem where a set of rectangular strips with fixed coordinates in one dimension enter a container, and their size is minimized by adjusting another dimension. If the height of the container represents the temporal domain of memory allocation, then the width of the container represents the memory usage.

[0073] Therefore, the advantage of this application lies in the fact that it proposes using the memory overlap of tensors as an important indicator for memory allocation. The memory overlap of tensors can serve as a reference and basis for measuring their influence on the memory allocation of other tensors, and can also generally reflect the degree of memory consumption during the lifetime of a tensor. Therefore, the offset calculation strategy based on the memory overlap of tensors can reduce the existence of memory fragmentation to a certain extent. Attached Figure Description

[0074] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to limit the scope of the invention.

[0075] Figure 1This is a schematic diagram of the topology of the computation flow graph.

[0076] Figure 2 This is a schematic diagram of the lifecycle and scheduling state of tensors 1 and 3.

[0077] Figure 3(a) is a schematic diagram of the lifecycle of all tensors.

[0078] Figure 3(b) is a schematic diagram of the degree of memory overlap.

[0079] Figure 4 This is a diagram illustrating the search for memory gaps.

[0080] Figure 5 This is a complete scheduling flowchart for memory allocation.

[0081] Figure 6 This is a flowchart illustrating the method described in this application. DETAILED DESCRIPTION

[0082] To better understand the technical content and advantages of the present invention, the present invention will now be described in further detail with reference to the accompanying drawings.

[0083] This invention belongs to the field of deep learning compilers and is applied to memory management optimizations in the compiler backend. Effectively reusing memory buffers can improve cache hit rate and inference speed.

[0084] The concepts and processes described below will be based on Figure 1 The intermediate tensors of the computation graph are described using the computation graph as an example. They have the following key characteristics:

[0085] Lifecycle: The lifecycle of a tensor represents the entire period during which the tensor is active in memory. The lifecycle of an intermediate tensor *t* can be defined as {start point, end point}, where the start point and end point are the indices of the producer operator of the intermediate tensor *t* and the last operator that used the intermediate tensor *t* as its input, respectively. These indices come from a topological ordering of the neural network, which is also the execution order of the operators. It is worth noting that no two tensors with intersecting intervals can share memory. Figure 2 The lifetime of tensor 1 is {1,2}, and the lifetime of tensor 3 is {3,5}. Memory scheduling state: The memory scheduling state represents the state of a tensor in memory throughout the entire inference timeline. The memory scheduling state of intermediate tensor t can be defined as {activity, tensor size}, where activity represents whether the tensor exists in memory, and tensor size is the amount of space occupied by the data in the tensor in memory (aligned size in bytes). Figure 2 If the inference timing is 2, the memory scheduling state of tensor 1 is {1,20}, and the lifetime of tensor 3 is {0,5}.

[0086] This method is based on an offset calculation strategy based on the degree of memory overlap of tensors, as shown in Figures 3(a) and 3(b):

[0087] A large block of memory is pre-allocated, and intermediate tensors are data buffers partitioned from the memory block using offsets. This method is called memory offset calculation, and its goal is to minimize the size of the allocated memory block. The offset calculation problem can be viewed as a special case of the two-dimensional stripe filling problem. This problem can be abstracted and simplified into a filling problem where a set of rectangular strips with fixed coordinates in one dimension are placed into a container, and their size is minimized by adjusting another dimension. If the height of the container represents the temporal domain of memory allocation, then the width of the container represents the memory usage.

[0088] The degree of memory overlap of a tensor can serve as a reference and basis for measuring its influence on the memory allocation of other tensors, and can also generally reflect the degree of memory consumption during the tensor's lifetime. Therefore, a strategy for calculating offsets based on the degree of memory overlap of tensors is explored, which can reduce the existence of memory gaps to some extent.

[0089] Therefore, as Figure 6 As shown, this application discloses a memory offset calculation method based on the degree of memory overlap of tensors. For static memory planning, the method includes the following steps:

[0090] S1, sort the allocation priority of intermediate tensors in the computation graph based on the degree of memory overlap of tensors;

[0091] The sorting of the memory overlap of the tensors further includes:

[0092] The rule for sorting tensors by memory overlap is to sort them in a non-increasing manner according to the following indicators: tensor overlap degree, tensor size, and starting point; the earlier the indicator appears, the higher its priority. The memory overlap degree is assumed to be as shown in the table below:

[0093]

[0094] The order of memory overlap of tensors is: 2->5->4->3->1->0->6;

[0095] S2, memory gap search, finding a suitable memory gap to match the corresponding intermediate tensor:

[0096] Memory gaps can be viewed as memory buffers that have been used before a certain point in time and are currently in a freed state; further including: for each intermediate tensor,

[0097] First, examine the allocated tensors whose lifetimes intersect with the current tensor to find the minimum memory gap between them, so that the current tensor fits the gap.

[0098] If such a gap is found, the current tensor is assigned to this gap;

[0099] Otherwise, after allocating its current tensor to the lowest tensor intersecting with its lifetime, the corresponding offset is assigned to the current tensor, and the tensor is in an allocated state; the search method for memory gaps is as follows: Figure 4 As shown.

[0100] In summary, the complete memory gap search process is further described as follows: When processing the intermediate tensor T of the memory to be allocated... a The following design steps are required:

[0101] First, it will be based on tensor T a By analyzing key information such as data type, data layout, and dimensions, the required memory size S can be accurately calculated. a This step is crucial, as it ensures an accurate understanding of memory requirements and lays a solid foundation for subsequent memory allocation.

[0102] Next, the list of allocated tensors will be carefully traversed to identify those intermediate tensors T that are related to the current tensor T to be allocated in terms of lifetime. a Tensor sets V with intersection t To manage these tensors more efficiently, they will be sorted in ascending order according to their starting addresses, resulting in a clearer and more orderly memory layout. Subsequently, the tensor set V will be used... t To retrieve the memory gap list V f By comparing the size of each memory gap, denoted as f0, f1, ..., f n It will search for the first tensor T whose size is greater than or equal to the tensor T. a The required memory and the most suitable memory gap; in this process, the best fit principle is followed, that is, min(f x -S a If a suitable memory gap V is successfully found, its value is greater than or equal to 0. x It will then be assigned to tensor T. a Make its starting address the same as V x The starting addresses are the same;

[0103] However, if a sufficiently large memory gap cannot be found to satisfy tensor T a If the requirement is met, the memory expansion function will be activated; at this time, the tensor T will be... a Assigned to the already assigned tensor set V t The last tensor T bAfter that, ensure that the starting address of tensor Ta immediately follows T. b Afterwards, and considering T b Required memory size;

[0104] Finally, to ensure the consistency and integrity of memory management, the intermediate tensor T of successfully allocated memory will be... a Add it to the list of assigned tensors for later management and querying;

[0105] By carefully designing and executing the above steps, we can ensure efficient memory utilization and reasonable allocation, thus providing a basis for tensor T. a This provides a strong guarantee for the smooth operation of the system.

[0106] S3, overall memory allocation and scheduling:

[0107] The offset calculation strategy based on the memory overlap of the tensors in steps S1 and S2 performs a complete scheduling of memory allocation for the neural network topology, as follows: Figure 5 As shown. Based on the offset calculation strategy of the memory overlap degree of the tensors in steps S1 and S2, a complete scheduling of memory allocation for the neural network topology is performed. The complete scheduling process of memory allocation is as follows:

[0108] S3.1 Resource Initialization:

[0109] 1. Construct the allocated list:

[0110] Before any new memory allocation begins, an allocated list needs to be maintained, recording all tensors that have been allocated memory in the current system and their related information. This list typically includes information such as the tensor's identifier, starting address, required memory size, and lifetime. The purpose of building this list is to ensure that when making new memory allocations, it is clear which memory regions are already occupied, thereby avoiding memory conflicts and overlaps.

[0111] 2. Construct a list of overlapping activity levels:

[0112] The activity overlap list is used to record allocated tensors whose lifecycles intersect with the tensor to be allocated. By constructing this list, tensors that may cause memory conflicts with the tensor to be allocated can be quickly located, thus enabling more accurate memory allocation decisions. The method for constructing this list is usually to traverse the allocated list, compare the lifecycle of each allocated tensor with the lifecycle of the tensor to be allocated, and find the tensors that have intersections.

[0113] S3.2, Allocation and Scheduling:

[0114] 1. Calculate the memory required for the tensor to be allocated:

[0115] Based on the data type, data arrangement, and dimensions of the tensor to be allocated, the required memory size can be calculated; this size will serve as the basic basis for memory allocation.

[0116] 2. Retrieve memory gaps:

[0117] After obtaining information about the memory size required for the tensor to be allocated, the next step is to search for available memory gaps in the allocated list. This is typically done by iterating through the tensors in the active overlap list, calculating the memory gap size between them, and finding the gap that meets the requirements of the tensor to be allocated.

[0118] 3. Select appropriate memory gaps:

[0119] After retrieving multiple available memory gaps, it is necessary to select the most suitable gap according to a certain strategy. This strategy includes the best fit principle, that is, selecting a gap whose size is closest to the memory required by the tensor to be allocated and leaving a certain margin, so as to make full use of memory resources and reduce fragmentation.

[0120] 4. Allocate memory:

[0121] Once a suitable memory gap is selected, the tensor to be allocated can be allocated into that gap; this involves updating the starting address and memory usage of the tensor to be allocated, and adding the corresponding record to the allocated list.

[0122] 5. Handling insufficient memory situations:

[0123] If a suitable memory gap cannot be found in the active overlap list to meet the needs of the tensor to be allocated, measures to expand the memory space are needed to handle the insufficient memory situation.

[0124] 6. Update the assigned list and the list of overlapping activity levels:

[0125] After memory allocation is complete, the allocated list and the active overlap list need to be updated to reflect the latest memory usage; this includes adding newly allocated tensors to the allocated list and updating the active overlap list as needed.

[0126] The above steps enable resource initialization and allocation scheduling, ensuring that tensor memory allocation is both efficient and accurate. This process can be adjusted and optimized according to specific system requirements and optimization goals to achieve better performance and resource utilization.

[0127] For example: allocating memory for tensors based on their degree of memory overlap. Figure 5 In step (1), memory is allocated for tensor 2. The allocated list is empty, so the activity overlap list is also empty. The offset starting point of tensor 2 is set to 0. After the allocation is completed, tensor 2 is added to the allocated list, and then... Figure 5 In step (2), memory is allocated for tensor 5, and the activity overlap list of the current tensor 5 is calculated. The activity overlap list only contains tensor 2, and there is no suitable memory gap. Therefore, the memory space is expanded to place tensor 5 immediately after tensor 1.

[0128] When allocating memory for tensors based on memory overlap, the following standard procedure should be followed:

[0129] Figure 5 (1) Allocate memory for tensor 2.

[0130] Calculate the memory required for Tensor 2: Based on the data type, size, and other information of Tensor 2, accurately calculate its required memory size.

[0131] Memory gap retrieval: The allocated list is empty, therefore the activity overlap list is also empty.

[0132] Allocate memory: Set the starting offset of tensor 2 to 0.

[0133] Update the allocated list: Add tensor 2, which has been allocated memory, to the allocated list and record its starting address, size, and other information.

[0134] Figure 5 (2) Allocate memory for tensor 5.

[0135] Calculate the memory required for tensor 5: Similarly, calculate the memory required for tensor 5 based on its relevant information. Calculate the activity overlap list: Traverse the allocated list, and based on the lifecycle of tensor 5, calculate the activity overlap with other allocated tensors, constructing an activity overlap list. In this example, the activity overlap list contains only tensor 2.

[0136] Searching for memory gaps: For tensor 2 in the active overlap list, check if there are suitable memory gaps around it. Since there are no suitable memory gaps in this case, it is necessary to expand the memory space: Since there are no suitable memory gaps, we choose to expand the memory space. Specifically, we allocate tensor 5 immediately after a tensor in the allocated list (such as tensor 1), ensuring that it does not conflict with other allocated tensors.

[0137] Update the allocated list: Add tensor 5, which has been allocated memory, to the allocated list and update its relevant information.

[0138] By following the above standardized procedures, we can ensure that the memory allocation process considers both the degree of memory overlap to improve memory utilization efficiency and the coordination and synchronization between various tensors. This helps reduce memory fragmentation, optimize memory layout, and improve the overall performance of the system.

[0139] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for calculating memory offset based on the degree of memory overlap using tensors, characterized in that, The method for static memory planning includes the following steps: S1, sort the allocation priority of intermediate tensors in the computation graph based on the degree of memory overlap of tensors; The sorting of the memory overlap of the tensors further includes: The rule for sorting the memory overlap of tensors is to sort them in a non-increasing manner according to the following indicators: tensor overlap, tensor size, and starting point. That is, when evaluating the memory overlap of tensors, there are three indicators: tensor overlap, tensor size, and starting point, and their importance decreases in that order. S2, memory gap search, finding a suitable memory gap to match the corresponding intermediate tensor; further includes: For each intermediate tensor, first check the allocated tensors whose lifetimes intersect with the current tensor to find the smallest memory gap between them, so that the current tensor fits the gap; if such a gap is found, the current tensor is allocated to this gap; otherwise, after allocating the current tensor to the bottommost tensor whose lifetime intersects with its own, the corresponding offset is assigned to the current tensor, and the tensor is in the allocated state. S3, overall memory allocation and scheduling: The offset calculation strategy based on the memory overlap of the tensors in steps S1 and S2 is used to perform a complete scheduling of memory allocation for the neural network topology. The complete scheduling process of memory allocation is as follows: S3.1, Resource initialization:

1. Construct the allocated list: Before any new memory allocation begins, an allocated list needs to be maintained, recording all tensors that have already been allocated memory in the current system and their related information; this list typically includes information such as the tensor's identifier, starting address, required memory size, and lifetime.

2. Construct a list of overlapping activity levels: The activity overlap list records allocated tensors whose lifecycles intersect with the tensor to be allocated. This list is typically constructed by iterating through the allocated list, comparing the lifecycle of each allocated tensor with the lifecycle of the tensor to be allocated, and identifying the tensors with intersections. S3.2, Allocation Scheduling:

1. Calculate the memory required for the tensor to be allocated: Based on the data type, data arrangement, and dimensions of the tensor to be allocated, the required memory size can be calculated; this size will serve as the basic basis for memory allocation.

2. Retrieve memory gaps: After obtaining information about the memory size required for the tensor to be allocated, the next step is to search for available memory gaps in the allocated list.

3. Select appropriate memory gaps: After retrieving multiple available memory gaps, it is necessary to select the most suitable gap according to a certain strategy. This strategy includes the best fit principle, that is, selecting a gap whose size is closest to the memory required by the tensor to be allocated and leaving a certain margin, so as to make full use of memory resources and reduce fragmentation.

4. Allocate memory: Once a suitable memory gap is selected, the tensor to be allocated can be allocated into that gap; this involves updating the starting address and memory usage of the tensor to be allocated, and adding the corresponding record to the allocated list.

5. Handling insufficient memory situations: If a suitable memory gap cannot be found in the active overlap list to meet the needs of the tensor to be allocated, measures to expand the memory space are needed to handle the insufficient memory situation.

6. Update the assigned list and the list of overlapping activity levels: After memory allocation is complete, the allocated list and the active overlap list need to be updated to reflect the latest memory usage.

2. The memory offset calculation method based on tensor-based memory overlap as described in claim 1, characterized in that, In step S1 First, the tensor overlap is the most critical indicator for evaluating the degree of memory overlap, directly reflecting the degree of overlap between different tensors in memory, that is, the size of the memory space they share. The higher the overlap, the higher the memory utilization efficiency, but it may also increase the complexity of memory management. Therefore, tensors with higher overlap should be given priority in the sorting process in order to better optimize memory utilization. Secondly, the tensor size is also an important indicator; it determines the amount of space a tensor occupies in memory. Larger tensors require more memory, so they should be taken into consideration when sorting. By considering the tensor size, we can better balance memory usage and performance requirements, and avoid situations where memory is insufficient due to an excessively large tensor. Finally, although the starting point is relatively minor, it can still affect the degree of memory overlap. The starting point determines the initial position of the tensor in memory, and different starting points may lead to different memory layouts and overlaps. Therefore, the influence of the starting point should also be considered appropriately during the sorting process to ensure the rationality of the memory layout. In summary, when evaluating the degree of tensor memory overlap, three indicators should be considered comprehensively: tensor overlap, tensor size, and starting point, and ranked in order of their importance.

3. The memory offset calculation method based on tensor-based memory overlap as described in claim 1, characterized in that, In step S1, the intermediate tensor of the computation graph has the following characteristics: Lifecycle: The lifecycle of a tensor represents the entire process during which the tensor is active in memory; Memory scheduling state: The memory scheduling state represents the state of the tensor in memory throughout the entire inference timeline; Tensor overlap: Tensor overlap represents the degree to which tensor memory states overlap during their active periods. Tensor overlap is a key indicator for measuring the degree to which different tensors share memory space, and it is influenced by two core factors: overlap size and the number of overlapping tensors. First, the overlap size directly reflects the total size of other tensors that intersect with it during the tensor's lifetime; the calculation of the overlap size usually involves a precise analysis of the tensor's layout and location in memory to determine the overlapping area between them. Secondly, the number of overlapping tensors is also an important factor affecting tensor overlap. It represents the number of active tensors that exist simultaneously in memory; when the number of overlapping tensors increases, it means that more tensors are active in the memory space, which helps to reduce memory fragmentation and waste; however, as the number of overlapping tensors increases, the complexity of memory management may also increase accordingly, requiring more refined strategies to ensure coordination and synchronization between the tensors. By taking into account both the overlap size and the number of overlap tensors, a more comprehensive assessment of tensor overlap can be achieved.

4. The memory offset calculation method based on tensor-based memory overlap as described in claim 3, characterized in that, The features of the intermediate tensor further include: The lifetime of an intermediate tensor t can be defined as {start point, end point}, where the start point and end point are the producer operator of the intermediate tensor t and the index of the last operator that took the intermediate tensor t as its input, respectively. These indices come from a topological sorting of the neural network, which is also the execution order of the operators. It is worth noting that no two tensors with intersecting intervals can share memory. The memory scheduling state of an intermediate tensor t can be defined as {activity, tensor size}, where activity represents whether the tensor exists in memory, and tensor size is the amount of space occupied by the data in the tensor in memory, in bytes.

5. The memory offset calculation method based on tensor memory overlap as described in claim 3, characterized in that, In step S1, the degree of memory overlap is assumed to be as shown in the following table: The order of memory overlap of tensors is: 2->5->4->3->1->0->6; During the tensor memory overlap sorting process, tensors are first sorted according to their overlap to ensure that tensors with smaller overlap are processed first. As shown in the table, Tensor 2 has an overlap size of 54 and 5 overlapping tensors, while Tensor 3 has an overlap size of 30 and 3 overlapping tensors. Therefore, Tensor 2 has an overlap of 0, and Tensor 3 has an overlap of 2. Consequently, after sorting, Tensor 2 comes before Tensor 3. Then, among tensors with the same overlap, tensors with larger sizes are processed first. As shown in the table, Tensor 3 has a tensor size of 5, and Tensor 4 has a tensor size of 7. Therefore, after sorting, Tensor 4 comes before Tensor 3. Finally, if the first two criteria are the same, tensors with earlier starting points are processed first.

6. The memory offset calculation method based on tensor-based memory overlap as described in claim 1, characterized in that, In step S2, during the search for memory gaps: Memory gaps can be viewed as memory buffers that have been used before a certain point in time and are currently in a freed state.

7. The memory offset calculation method based on tensor-based memory overlap as described in claim 1, characterized in that, In step S2, the complete memory gap search process is described as follows: When processing the intermediate tensor T of the memory to be allocated... a The design process involves the following steps: First, based on tensor T... a By analyzing key information such as data type, data layout, and dimensions, the required memory size S can be accurately calculated. a ; Next, the list of allocated tensors will be carefully traversed to identify those intermediate tensors T that are related to the current tensor T to be allocated in terms of lifetime. a Tensor sets V with intersection t To manage these tensors more efficiently, they will be sorted in ascending order according to their starting addresses, resulting in a clearer and more orderly memory layout. Subsequently, the tensor set V will be used... t To retrieve the memory gap list V f By comparing the size of each memory gap, denoted as f0, f1, ..., f n It will search for the first tensor T whose size is greater than or equal to the tensor T. a The required memory and the most suitable memory gap; in this process, the best fit principle is followed, that is, min(f x -S a If a suitable memory gap V is successfully found, its value is greater than or equal to 0. x It will then be assigned to tensor T. a Make its starting address the same as V x The starting addresses are the same; However, if a sufficiently large memory gap cannot be found to satisfy tensor T a If the requirement is met, the memory expansion function will be activated; at this time, the tensor T will be... a Assigned to the already assigned tensor set V t The last tensor T b After that, ensure that the starting address of tensor Ta immediately follows T. b Afterwards, and considering T b Required memory size; Finally, to ensure the consistency and integrity of memory management, the intermediate tensor T of successfully allocated memory will be... a Add it to the list of assigned tensors for later management and querying; By carefully designing and executing the above steps, we can ensure efficient memory utilization and reasonable allocation, thus providing a basis for tensor T. a This provides a strong guarantee for the smooth operation of the system.

8. The memory offset calculation method based on tensor-based memory overlap as described in claim 1, characterized in that, In step S3, the offset calculation strategy based on the memory overlap of tensors includes: A large block of memory is pre-allocated, and the intermediate tensor is a data buffer partitioned by offsets within the memory block. This method is called memory offset calculation, and its goal is to minimize the size of the allocated memory block. The offset calculation problem can be viewed as a special case of the two-dimensional strip filling problem. This problem can be abstracted and simplified into a filling problem in which a set of rectangular strips with fixed coordinates in a certain dimension enter a container and their size is minimized by adjusting another dimension. If the height of the container represents the time domain of memory allocation, then the width of the container represents the memory usage.