Memory management method, storage medium and electronic equipment

By optimizing the memory allocation strategy based on the execution order and memory demand sequence of the network computation graph in static memory management, the memory fragmentation problem is solved, and memory utilization and management efficiency are improved.

CN120909952APending Publication Date: 2025-11-07MONTAGE TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410559189.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing static memory management methods are prone to memory fragmentation during use, leading to reduced memory utilization, especially in performance-critical applications where it is difficult to effectively utilize limited memory space.

Method used

Based on the execution order of each operator in the network computation graph, the data flow sequence and memory requirement sequence are obtained. By constructing target memory blocks, memory space is allocated to each memory requirement according to the order and lifecycle of the memory requirements. Memory allocation is optimized by strategies such as memory space reuse, reordering, relocation, and splitting operators.

Benefits of technology

It reduces memory fragmentation, improves memory utilization, reduces communication overhead, and improves memory management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909952A_ABST
    Figure CN120909952A_ABST
Patent Text Reader

Abstract

The invention provides a memory management method, a storage medium and electronic equipment. The memory management method comprises the following steps: according to an execution sequence of each operator in a network calculation graph, obtaining a data flow sequence corresponding to each operator during operation, and constructing a memory demand sequence corresponding to the data flow sequence; creating a target memory block in the available memory space of the memory; and allocating a memory space for each memory demand according to the size of the target memory block and the sequence and life cycle of each memory demand in the memory demand sequence. The memory management method is favorable for improving the memory utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computers, and relates to a memory management method, in particular to a memory management method, a storage medium and an electronic device. BACKGROUND

[0002] Memory is an important component of electronic products. Efficient management and optimization of the storage space of the memory can improve the utilization rate of the storage space, thereby achieving the effect of reducing cost and improving hardware running performance. For chips that need to calculate a large amount of data (such as AI chips), how to efficiently and reasonably use the limited memory space is particularly important.

[0003] Memory management generally includes static memory management and dynamic memory management. Static memory management allocates memory at the compilation stage, and is more advantageous in some applications with relatively high performance requirements and relatively stable memory requirements. However, the existing static memory management method usually produces memory fragmentation during use, reducing memory utilization.

[0004] Therefore, there is a need for a memory management method that can improve memory utilization. SUMMARY

[0005] Embodiments of the present application provide a memory management method, a storage medium and an electronic device, which can reduce the generation of memory fragmentation and improve memory utilization.

[0006] In a first aspect, an embodiment of the present application provides a memory management method, which includes: obtaining a data flow sequence corresponding to the runtime of each operator according to the execution order of each operator in a network computation graph, and constructing a memory demand sequence corresponding to the data flow sequence; creating a target memory block in the available memory space of a memory; and allocating memory space for each memory demand according to the size of the target memory block and the order and life cycle of each memory demand in the memory demand sequence.

[0007] In an implementation form of the first aspect, obtaining a data flow sequence corresponding to the runtime of each operator according to the execution order of each operator in a network computation graph, and constructing a memory demand sequence corresponding to the data flow sequence includes: obtaining a data flow sequence corresponding to each operator according to the execution order of each operator and the input data, parameters and output data corresponding to the runtime of each operator; determining the memory demand of each data according to the size of each data in the data flow sequence, and forming the memory demand sequence according to the memory demand of each data.

[0008] In an implementation form of the first aspect, the order of each memory demand in the memory demand sequence is determined according to the order of the data corresponding to each memory demand in the data flow sequence.

[0009] In an implementation form of the first aspect, the allocating memory space for each memory demand in the memory demand sequence according to the size of the target memory block and the order and the life cycle of each memory demand in the memory demand sequence comprises: determining and recording the first M memory demands in the memory demand sequence that can be simultaneously satisfied by the target memory block according to the order of each memory demand in the memory demand sequence, where M is an integer less than or equal to N, and N is the number of memory demands in the memory demand sequence; if M is not 0, obtaining the life cycle of each of the first M memory demands, and allocating memory space for the first M memory demands from the target memory block according to the life cycle.

[0010] In an implementation form of the first aspect, the determining the first M memory demands in the memory demand sequence that can be simultaneously satisfied by the target memory block comprises: step a, judging whether the size of the memory space required by the kth memory demand in the memory demand sequence memory_k_size is less than or equal to a reference comparison value block_ref_size, where k is a positive integer, the initial value of k is 1, and the initial value of the reference comparison value is equal to the size of the memory space of the target memory block; if memory_k_size is less than or equal to block_ref_size, entering step b; if memory_k_size is greater than block_ref_size, entering step d; step b, recording the kth memory demand, and updating the reference comparison value as block_ref_size-memory_k_size; step c, making the value of k plus 1, and judging whether the value of k is greater than the number N of memory demands in the memory demand sequence; if the value of k is greater than N, ending the flow; if the value of k is less than or equal to N, returning to step a until the value of k is greater than N or memory_k_size is greater than block_ref_size; step d, determining M as (k-1).

[0011] In an implementation form of the first aspect, the obtaining the life cycle of each of the first M memory demands comprises: obtaining the life cycle of each of the first M memory demands according to the dependency relationship between the data corresponding to each memory demand and the order of each memory demand.

[0012] In an implementation form of the first aspect, the allocating memory space for each memory demand in the memory demand sequence according to the size of the target memory block and the order and the life cycle of each memory demand in the memory demand sequence comprises: determining and recording the first M memory demands in the memory demand sequence that can be simultaneously satisfied by the target memory block according to the order of each memory demand in the memory demand sequence, where M is an integer less than or equal to N, and N is the number of memory demands in the memory demand sequence; if M is not 0, obtaining the life cycle of each of the first M memory demands, and allocating memory space for the first M memory demands from the target memory block according to the life cycle.

[0013] In an implementation form of the first aspect, after allocating memory space for the first M memory demands from the target memory block according to the life cycle, the memory management method further comprises: if there are still memory demands in the memory demand sequence that need to be allocated memory space, or a new memory demand sequence is generated, re-creating the target memory block.

[0014] In an implementation form of the first aspect, if the M is 0, the memory management method further comprises: allocating memory space for the first memory demand in the memory demand sequence according to a preset strategy.

[0015] In an implementation form of the first aspect, the preset strategy comprises any one or a combination of more than one of: memory space reuse, reordering allocated memory space, moving and releasing at least part of the allocated memory, splitting an operator.

[0016] In an implementation form of the first aspect, allocating memory space for the first memory demand according to the memory space reuse strategy comprises: searching for a first operator in an operator corresponding to the first memory demand, wherein the first memory demand corresponds to output data of the first operator, and input data of the first operator has been allocated memory space; when the first operator is found, reusing the memory space allocated for the input data of the first operator for the first memory demand.

[0017] In an implementation form of the first aspect, allocating memory space for the first memory demand according to the reordering allocated memory space strategy comprises: judging whether the sum of the available memory space of the memory can meet the first memory demand; if yes, arranging each allocated memory in turn from the start address or the end address of the memory by a memory moving module, and making each allocated memory space continuous; and after the arrangement, allocating memory space for the first memory demand from the remaining available memory space of the memory.

[0018] In an implementation form of the first aspect, allocating memory space for the first memory requirement according to the strategy of moving and releasing at least part of the allocated memory includes: searching for a second operator in the operators corresponding to the first memory requirement, wherein the second operator is an operator whose first state is an undefined state among the operators corresponding to the first memory requirement, and the undefined state is used to indicate that all memory requirements corresponding to the operator are not allocated memory space; searching for memory requirements not corresponding to the second operator from the memory requirements corresponding to the allocated memory space and recording; selecting at least part of the memory requirements from the recorded memory requirements, and moving storage data in memory space corresponding to the selected at least part of the memory requirements to another memory to release the memory space corresponding to the selected at least part of the memory requirements, wherein the continuous available memory space formed by the memory space corresponding to the selected at least part of the memory requirements and the available memory space in the memory is greater than or equal to the memory space required by the first memory requirement; and allocating memory space for the first memory requirement from the continuous available memory space.

[0019] In an implementation form of the first aspect, allocating memory space for the first memory requirement according to the strategy of splitting the operator includes: searching for a third operator in the operators corresponding to the first memory requirement and splitting the third operator, wherein the first memory requirement corresponds to other data of the third operator except input data; after searching for the third operator and splitting the third operator, if the available memory space of the memory can meet the first memory requirement, reconstructing the memory requirement sequence and allocating memory space for the first memory requirement.

[0020] In an implementation form of the first aspect, the network computing graph is a directed acyclic graph.

[0021] In an implementation form of the first aspect, the target memory block is a continuous memory space in the available memory space of the memory.

[0022] In the second aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the memory management method in any one of the embodiments of the first aspect of the present application.

[0023] In the third aspect, the embodiments of the present application provide an electronic device, which includes a memory storing a computer program, and a processor connected to the memory in communication, and the processor executes the memory management method in any one of the embodiments of the first aspect of the present application when invoking the computer program.

[0024] The memory management method provided by the embodiments of the present application allocates memory spaces for each memory requirement according to the size of the target memory block and the order and life cycle of each memory requirement in the memory requirement sequence. In this way, the generation of memory fragmentation can be reduced, and the memory utilization rate can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 A schematic diagram of an application scenario is shown as an embodiment of the present application.

[0026] Figure 2A A flowchart of the memory management method provided by the embodiments of the present application is shown.

[0027] Figure 2B A detailed flowchart of step S21 in the embodiments of the present application is shown.

[0028] Figure 2C An example diagram of a network computing graph is shown as an embodiment of the present application.

[0029] Figure 3 A detailed flowchart of step S23 in the embodiments of the present application is shown.

[0030] Figure 4 A schematic diagram of memory allocation and release is shown as an embodiment of the present application.

[0031] Figure 5 A schematic diagram of memory reordering is shown as an embodiment of the present application.

[0032] Figure 6 A flowchart of allocating a memory space for the first memory requirement according to the strategy of moving and releasing at least part of the allocated memory is shown as an embodiment of the present application. DETAILED DESCRIPTION

[0033] The embodiments of the present application are described below through specific and detailed examples. Those skilled in the art can easily understand other advantages and effects of the present application from the disclosure of the present specification. The present application can also be implemented or applied through other different specific embodiments, and each detail in the present specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0034] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present application, and only the components related to the present application are shown in the diagrams, not the number, shape and size of the components when actually implemented. The shapes, number and proportions of each component in actual implementation can be arbitrarily changed, and the layout pattern of the components can also be more complex.

[0035] Generally, in the design and optimization process of chips that need to calculate a large amount of data, efficiently managing the memory resources of the memory (such as DDR and SRAM) is an effective way to reduce costs. Taking an AI chip as an example, efficient management of SRAM can enhance the reuse of internal memory of the AI chip, thereby effectively reducing the bandwidth of network inference. The bandwidth of network inference is reduced, which on the one hand can improve the stability of model inference time and reduce the problem of inference time decline caused by bandwidth competition; on the other hand, it is beneficial to improve the bandwidth bottleneck and the execution speed of operators. However, the memory management algorithm used in the prior art solution usually produces memory fragmentation in use, increases communication overhead, and is difficult to achieve the best effect. For example, exchanging unused data from the memory space of the current operation device to other memory space, and exchanging the data from other memory space back when accessed next time, will generate communication and synchronization overhead.

[0036] At least for the above problems, the embodiments of the present application provide a memory management method. The memory management method can be applied to an electronic device, for example, an electronic device including an AI chip, but the embodiments of the present application do not limit the type of electronic device. For example, the memory management method can also be applied to a smart terminal, an embedded system, a cloud server, etc.

[0037] Figure 1 A structural schematic diagram of an electronic device is shown. As shown in Figure 1 The electronic device includes at least one processor and a memory. The processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a microprocessor, and an application specific integrated circuit (ASIC), etc. The processor is configured to execute various types of instructions and operations, for example, to execute software or firmware programs stored in the memory, so that the electronic device provides a variety of functions and services. For example, the processor can execute programs or process data to implement the static memory management method provided by the embodiments of the present application.

[0038] The memory can be all types of memory suitable for static memory management, such as DDR, SRAM, etc. The memory can be used to store program instructions and data for the processor to call to implement the static memory management method provided by the embodiments of the present application.

[0039] In some embodiments, the electronic device can be an AI device, which includes an AI chip. The processor can include a first processor and a second processor, where the first processor can be a processor outside the AI chip, such as a CPU, and the second processor can be a processor inside the AI chip and can include a register, such as an NPU. The memory can include a first memory and a second memory, where the first memory can be, for example, a DDR, which can be located outside the AI chip, and the second memory can be, for example, a SRAM, which can be located inside the AI chip.

[0040] The first processor can be responsible for processing various general computing tasks in the electronic device, including but not limited to running of an operating system, multitasking, system management, general data processing, etc.

[0041] In some embodiments, the first processor can include a driver of the second processor, and the first processor can configure the second processor through the driver. For example, the first processor can configure the second processor to process specified data and allocate registers of the second processor through the driver. For example, in some application scenarios, each frame of image captured by a camera can be automatically stored in a certain memory, and each time an image is stored, the first processor can issue an execution command to the second processor, instructing the second processor to call the image from the memory for AI model inference.

[0042] In some embodiments, the second processor can be a neural network computing processor, which can quickly process input information by drawing on the structure of a biological neural network, such as the transmission mode between human brain neurons, and can also constantly self-learn. Through the second processor, intelligent cognition applications of intelligent terminals can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc. The network code and parameters required by the second processor in the data processing process can be stored in the second memory.

[0043] For convenience of description, the memory management method provided by the embodiments of the present application will be described in detail below in combination with the accompanying drawings of the embodiments of the present application and by taking an AI device and a neural network model as examples. However, those skilled in the art can understand that the memory management method of the present application is applicable to all devices or systems that can adopt a static memory management method.

[0044] Figure 2A A flowchart of the memory management method provided by the embodiments of the present application is shown. As shown in Figure 2A The memory management method provided by the embodiments of the present application includes the following steps.

[0045] S21, according to the execution order of each operator in the network computing graph, obtaining a data stream sequence corresponding to the runtime of each operator, and constructing a memory requirement sequence corresponding to the data stream sequence.

[0046] In some embodiments, the network computing graph is a directed acyclic graph, which is composed of operators and edges. Each operator corresponds to an operation, such as splitting, convolution, softmax, etc., and the edges correspond to the data transmission (data input / output) relationship between the operators. Before performing the operations on the operators, the operators need to be sorted to determine the execution order of the operators. Any suitable method can be used to determine the execution order of the operators, which will not be described herein.

[0047] Referring to Figure 2B In some embodiments, according to the execution order of the operators in the network computing graph, the data flow sequence corresponding to the runtime of each operator is obtained, and the memory requirement sequence corresponding to the data flow sequence is constructed, which can include:

[0048] S211, according to the execution order of the operators and the input data, parameters and output data corresponding to the runtime of each operator, the data flow sequence corresponding to each operator is obtained.

[0049] For example, it is assumed that Figure 2C The execution order of the operators in the network computing graph is: operator A -> operator B -> operator C. The operator A needs input data E1 and parameter W1 at runtime, and outputs data E2; the operator B needs input data E1 and parameter W2 at runtime, and outputs data E3; the operator C needs input data E2, E3 and parameter W3 at runtime, and outputs data E4. According to the execution order of the operators, the data flow sequence: data E1, parameter W1, data E2, parameter W2, data E3, parameter W3, data E4 is obtained.

[0050] S212, according to the size of each data in the data flow sequence, the memory requirement of each data is determined, and the memory requirement sequence is formed according to the memory requirement of each data. The order of each memory requirement in the memory requirement sequence can be determined according to the order of the data corresponding to each memory requirement in the data flow sequence.

[0051] For example, continuing to refer to Figure 2C According to the data E1, parameter W1, data E2, parameter W2, data E3, parameter W3, data E4 included in the data flow sequence, 7 memory requirements are generated, which constitute the memory requirement sequence. The memory requirement corresponding to data E1 can be ranked first in the memory requirement sequence, the memory requirement corresponding to parameter W1 can be ranked second, and so on, and the memory requirement corresponding to data E4 can be ranked seventh.

[0052] In some embodiments, in order to distinguish the order of each memory requirement in the memory requirement sequence, each memory requirement can be marked according to the order of each memory requirement. For example, the memory requirements corresponding to data E1, parameter W1, data E2, parameter W2, data E3, parameter W3, data E4 in the data flow sequence can be marked as memory 0, memory 1, memory 2, memory 3, memory 4, memory 5, memory 6 respectively.

[0053] S22, creating a target memory block in the available memory space of the memory.

[0054] The target memory block (hereinafter also referred to as Memory Block) can be a continuous memory space in the available memory space of the memory of the electronic device. In order to enable the target memory block to meet more memory requirements, in some embodiments, the target memory block can be the largest continuous memory space in the available memory space of the memory.

[0055] It should be noted that the present embodiment is described by taking the example of performing step S21 first and then performing step S22. However, in actual applications, the present embodiment does not limit the execution order of steps S21 and S22. For example, step S22 can be performed first, and then step S21 can be performed.

[0056] S23, allocating memory space for each memory requirement according to the size of the target memory block and the order and life cycle of each memory requirement in the memory requirement sequence. In some embodiments, the life cycle of the memory requirement can be obtained by simulation.

[0057] Please refer to Figure 3 , specifically, step S23 further comprises:

[0058] S231, determining and recording the first M memory requirements in the memory requirement sequence that can be met by the target memory block at the same time according to the order of each memory requirement in the memory requirement sequence, wherein M is an integer less than or equal to N, and N is the number of memory requirements in the memory requirement sequence.

[0059] In some embodiments, the first M memory requirements can be determined by the following steps.

[0060] Step a, judging whether the size of the memory space required by the kth memory requirement in the memory requirement sequence memory_k_size is less than or equal to the reference comparison value block_ref_size, wherein k is a positive integer, the initial value of k is 1, and the initial value of the reference comparison value is equal to the size of the memory space of the target memory block.

[0061] If the memory_k_size is less than or equal to the block_ref_size, step b can be entered; if the memory_k_size is greater than the block_ref_size, it indicates that the available memory space of the target memory block cannot meet the kth memory requirement, and step d is entered.

[0062] In step b, the kth memory requirement is recorded, and the reference comparison value is updated as block_ref_size - memory_k_size.

[0063] In step c, the value of k is increased by 1, and it is determined whether the value of k is greater than the number N of memory requirements in the memory requirement sequence. If the value of k is greater than N, it indicates that all the memory requirements have been recorded before, and the process can be ended. If the value of k is less than or equal to N, step a is returned until the value of k is greater than N or the memory_k_size is greater than the block_ref_size.

[0064] In step d, M is determined as (k-1). It can be understood that all the memory requirements recorded in step b constitute the first M memory requirements.

[0065] In step d, when M is 0, it indicates that the target memory block cannot meet the first memory requirement in the memory requirement sequence, i.e., the memory requirement memory 0, which means that the memory space of the target memory block is too small and needs to be processed according to a preset strategy. Details can be referred to below.

[0066] In step d, when M is 0, it indicates that the target memory block cannot meet the first memory requirement in the memory requirement sequence, i.e., the memory requirement memory 0, which means that the memory space of the target memory block is too small and needs to be processed according to a preset strategy. Details can be referred to below.

[0067] In step S233, the lifetimes of the first M memory requirements are obtained, and memory spaces are allocated for the first M memory requirements from the target memory block according to the lifetimes.

[0068] In some embodiments, the lifetimes of the first M memory requirements can be obtained according to the dependency relationship between the data corresponding to each memory requirement and the order of the memory requirements.

[0069] Specifically, each operator can correspond to multiple memory requirements. Taking operator A as an example, data E1, parameter W1 and data E2 of operator A correspond to memory requirements memory 0, memory 1 and memory 2 respectively. Similarly, each memory requirement can also correspond to multiple operators. For example, data E1 corresponds to both operator A and operator B, and thus memory 0 corresponding to data E1 corresponds to both operator A and operator B. For any memory requirement, when all operators corresponding to the memory requirement are in the released state, the life cycle of the memory requirement is completed. For any operator, when all memory requirements corresponding to the operator are allocated with corresponding memory spaces, the operator is in the released state. Accordingly, the embodiments of the present application can obtain the life cycle of each memory requirement according to the data dependency relationship among the data in each memory requirement and the order of each memory requirement.

[0070] In some embodiments, when allocating memory spaces for the first M memory requirements according to the life cycle of each of the first M memory requirements, the memory spaces can be allocated to the first M memory requirements in the order of descending life cycle along the direction from one boundary of the target memory block to another boundary of the target memory block.

[0071] For example, the address range of the target memory block is [addr_a, addr_b], and addr_a and addr_b are addresses corresponding to two boundaries of the target memory block respectively. The memory spaces in the target memory block can be allocated to the memory requirements in the order of descending life cycle along the direction from addr_a to addr_b or the direction from addr_b to addr_a.

[0072] For example, the address range of the target memory block is [addr_a, addr_b], and addr_a and addr_b are addresses corresponding to two boundaries of the target memory block respectively. The memory spaces in the target memory block can be allocated to the memory requirements in the order of descending life cycle along the direction from addr_a to addr_b or the direction from addr_b to addr_a. Figure 4 For example, the address range of the target memory block is [addr_a, addr_b], and addr_a and addr_b are addresses corresponding to two boundaries of the target memory block respectively. The memory spaces in the target memory block can be allocated to the memory requirements in the order of descending life cycle along the direction from addr_a to addr_b or the direction from addr_b to addr_a. Figure 4 For example, the address range of the target memory block is [addr_a, addr_b], and addr_a and addr_b are addresses corresponding to two boundaries of the target memory block respectively. The memory spaces in the target memory block can be allocated to the memory requirements in the order of descending life cycle along the direction from addr_a to addr_b or the direction from addr_b to addr_a. Figure 4 For example, the address range of the target memory block is [addr_a, addr_b], and addr_a and addr_b are addresses corresponding to two boundaries of the target memory block respectively. The memory spaces in the target memory block can be allocated to the memory requirements in the order of descending life cycle along the direction from addr_a to addr_b or the direction from addr_b to addr_a.

[0073] After step S233, if there still exists a memory requirement in the memory requirement sequence for which memory space is to be allocated, or a new memory requirement sequence is generated, the target memory block can be re-created. For example, the memory space [addr_c, addr_b] released in step S232 is taken as a new target memory block. Figure 4

[0074] In step S234, memory space is allocated for the kth memory requirement according to a preset strategy. It can be understood that k here is equal to 1, i.e. the memory requirement memory 0 in the first position in the memory requirement sequence.

[0075] The preset strategy includes, but is not limited to, any one or any combination of the following: memory space reuse, reordering of allocated memory space, moving and releasing at least part of the allocated memory, splitting operator.

[0076] 1. Memory space reuse, which is applicable to an operator whose output memory can reuse its input memory and does not affect the operation result.

[0077] Specifically, a first operator is first searched for in the operator corresponding to the kth memory requirement, where the output data of the first operator corresponds to the kth memory requirement, and the input data of the first operator has been allocated memory space. When the first operator is found, the kth memory requirement can reuse the memory space allocated for the input data of the first operator. After the reused memory space is released, the target memory block can be re-created, and the value of k is increased by 1.

[0078] 2. Reordering of the allocated memory in the memory.

[0079] Specifically, it is first determined whether the sum of the available memory space of the memory can satisfy the kth memory requirement. If yes, the allocated memories are arranged in sequence from the start address (or end address) of the memory by the memory moving module, so that the allocated memory spaces are arranged continuously, thereby increasing the maximum continuous available memory space of the memory. Then, memory space is allocated for the kth memory requirement from the maximum continuous available memory space. After the memory space of the kth memory requirement is released, the target memory block can be re-created, and the value of k is increased by 1.

[0080] Figure 5 An example diagram showing reordering of the allocated memory is shown. As shown in FIG. 4, the memory space allocated for the input data of the first operator is moved to the end of the memory, and the memory space allocated for the output data of the first operator is moved to the start of the memory. Figure 5 ​As shown, before the reordering, the memory between [addr_d, addr_e] in the device memory is allocated to memory requirements 3, 6 and 8, and the memory allocated to memory requirements 3, 6 and 8 is not continuous. After the reordering, the memory allocated to memory requirements 3, 6 and 8 is continuous, thus the maximum continuous available memory space between [addr_d, addr_e] is increased, thus the size of the target memory block is increased.

[0081] 3. Moving and releasing at least part of the allocated memory.

[0082] Referring to Figure 6 In the embodiment, the allocation of memory space to the kth memory requirement according to the strategy of moving and releasing at least part of the allocated memory includes the following steps.

[0083] S61. Finding a second operator in the operators corresponding to the kth memory requirement, wherein the second operator is the operator whose state corresponding to the kth memory requirement is UNDEFINED (i.e. undefined state).

[0084] It can be understood that each operator can correspond to multiple memory requirements, and each operator corresponds to a state at each memory requirement corresponding to the operator, for example, before all the memory requirements corresponding to the operator are allocated memory space, the state of the operator at all the memory requirements corresponding to the operator is UNDEFINED; when part of the memory requirements corresponding to the operator are allocated memory space, the state of the operator at the part of the memory requirements corresponding to the operator is RESIDENT (i.e. resident state), and when all the memory requirements corresponding to the operator are allocated memory, the state of the operator at all the memory requirements corresponding to the operator is RELEASE (i.e. release state).

[0085] S62. Finding and recording the memory requirements corresponding to the second operator from the memory requirements corresponding to the allocated memory space.

[0086] S63. Selecting at least part of the memory requirements from the recorded memory requirements, and moving the storage data in the memory space corresponding to the selected at least part of the memory requirements to another memory to release the memory space corresponding to the selected at least part of the memory requirements. Wherein the continuous available memory space formed by the memory space corresponding to the selected at least part of the memory requirements and the available memory space in the memory is greater than or equal to the memory space required by the kth memory requirement.

[0087] In some embodiments, when the at least part of the memory requirements is selected from the recorded memory requirements, the at least part of the memory requirements with the least number of moving times can be selected.

[0088] It can be understood that if the subsequent storage data needs to be moved to another memory, the storage space of the part of the storage data can be reallocated.

[0089] S64, allocating memory space for the kth memory requirement from the continuous available memory space formed by the memory space corresponding to the at least part of the memory requirements and the available memory space in the memory. After the memory space of the kth memory requirement is released, the target memory block can be reconstructed, and the value of k is increased by 1.

[0090] 4. Splitting the operator.

[0091] The third operator corresponding to the kth memory requirement is searched and split, wherein the kth memory requirement corresponds to other data of the third operator except input data, such as output data or parameters. If the third operator is not found, the first operator corresponding to the kth memory requirement is split. After splitting, the memory space required by the kth memory requirement is reduced, and at this time, if the available memory space of the memory can meet the memory requirement of the kth memory requirement, the memory requirement sequence needs to be reconstructed before the memory space of the kth memory requirement is allocated, because the network computation graph is changed due to the splitting of the operator.

[0092] In some embodiments, before the strategy of splitting the operator is implemented, the strategy of moving and releasing at least part of the allocated memory and the strategy of reordering the allocated memory in the memory can be implemented first, and at this time, the available memory space of the memory has been maximized, and then the strategy of splitting the operator can be implemented.

[0093] If all the above strategies and strategy combinations cannot meet the memory space required by the kth memory requirement, it indicates that the available memory space of the memory is insufficient to meet the running requirement.

[0094] According to the above description, it can be known that the memory management method provided by the embodiments of the present application can reduce the data transmission amount between the on-chip and off-chip, reduce the memory fragmentation, and is beneficial to increase the utilization rate of the on-chip memory.

[0095] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the memory management method provided by the embodiments of the present application. Those skilled in the art can understand that all or part of the steps of the method described in the above embodiments can be instructed by a program to complete the processor. The program can be stored in a computer readable storage medium. The storage medium is a non-transitory medium, for example, a random access memory, a read only memory, a flash memory, a hard disk, a solid state disk, a magnetic tape, a floppy disk, an optical disc and any combination thereof. The storage medium can be any available medium which can be accessed by a computer or a data storage device such as a server, a data center and the like which integrates one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)) or a semiconductor medium (for example, a solid state disk (SSD)) and the like.

[0096] The embodiments of the present application further provide an electronic device, which comprises a memory and a processor. The memory stores a computer program. The processor is connected to the memory in communication and executes the memory management method provided by the embodiments of the present application when the computer program is called.

[0097] The above embodiments only illustrate the principles and effects of the present application, but are not used to limit the present application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical idea of the present application should be covered by the claims of the present application.

Claims

1. A memory management method characterized by comprising: The method comprises the following steps: According to the execution order of each operator in the network computing graph, the data stream sequence corresponding to the runtime of each operator is obtained, and a memory requirement sequence corresponding to the data stream sequence is constructed; A target memory block is created in the available memory space of the memory; According to the size of the target memory block and the order and life cycle of each memory requirement in the memory requirement sequence, memory space is allocated for each memory requirement.

2. The memory management method of claim 1, wherein, According to the execution order of each operator in the network computing graph, the data stream sequence corresponding to the runtime of each operator is obtained, and a memory requirement sequence corresponding to the data stream sequence is constructed, which comprises: According to the execution order of each operator and the input data, parameters and output data corresponding to the runtime of each operator, the data stream sequence corresponding to each operator is obtained; According to the size of each data in the data stream sequence, the memory requirement of each data is determined, and the memory requirement sequence is formed according to the memory requirement of each data.

3. The memory management method of claim 1, wherein, The order of each memory requirement in the memory requirement sequence is determined according to the order of the data corresponding to the memory requirement in the data stream sequence.

4. The memory management method of claim 1, wherein, According to the size of the target memory block and the order and life cycle of each memory requirement in the memory requirement sequence, memory space is allocated for each memory requirement, which comprises: According to the order of each memory requirement in the memory requirement sequence, the first M memory requirements in the memory requirement sequence that can be simultaneously satisfied by the target memory block are determined and recorded, wherein M is an integer less than or equal to N, and N is the number of memory requirements in the memory requirement sequence; If M is not 0, the life cycle of each of the first M memory requirements is obtained, and memory space is allocated for the first M memory requirements from the target memory block according to the life cycle.

5. The memory management method of claim 4, wherein, Determining the first M memory requirements in the memory requirement sequence that can be simultaneously satisfied by the target memory block comprises: Step a: determining whether the size of the memory space required by the kth memory requirement in the memory requirement sequence memory_k_size is less than or equal to the reference comparison value block_ref_size, wherein k is a positive integer, the initial value of k is 1, and the initial value of the reference comparison value is equal to the size of the memory space of the target memory block; If memory_k_size is less than or equal to block_ref_size, go to step b; if memory_k_size is greater than block_ref_size, go to step d; Step b: recording the kth memory requirement, and updating the reference comparison value to block_ref_size-memory_k_size; Step c: increasing the value of k by 1, and determining whether the value of k is greater than the number N of memory requirements in the memory requirement sequence; If the value of k is greater than N, the process is ended; if the value of k is less than or equal to N, return to step a until the value of k is greater than N or memory_k_size is greater than block_ref_size; Step d: determining M as (k-1).

6. The memory management method of claim 4, wherein, Obtaining the life cycle of each of the first M memory requirements comprises: According to the dependency relationship between data corresponding to each memory requirement and the order of each memory requirement, a life cycle of each of the first M memory requirements is obtained.

7. The memory management method of claim 4, wherein, Allocating memory space for the first M memory requirements from the target memory block according to the life cycles comprises: According to the order of the life cycles of the first M memory requirements from large to small, memory space is allocated for the first M memory requirements along the direction from one boundary of the target memory block to another boundary.

8. The memory management method of claim 4, wherein, After allocating memory space for the first M memory requirements from the target memory block according to the life cycles, the memory management method further comprises: If there are still memory requirements in the memory requirement sequence that need to be allocated memory space, or a new memory requirement sequence is generated, the target memory block is re-created.

9. The memory management method of claim 4, wherein, If the M is 0, the memory management method further comprises: allocating memory space for the first memory requirement in the memory requirement sequence according to a preset strategy.

10. The memory management method of claim 9, wherein, The preset strategy comprises any one or a combination of more than one of the following: memory space reuse, reordering of allocated memory space, moving and releasing at least part of the allocated memory, splitting operator.

11. The memory management method of claim 10, wherein, According to the strategy of memory space reuse, the memory space for the first memory requirement comprises: Finding a first operator in the operator corresponding to the first memory requirement, wherein the first memory requirement corresponds to the output data of the first operator, and the input data of the first operator has been allocated memory space; When the first operator is found, the first memory requirement is reused to allocate memory space for the input data of the first operator.

12. The memory management method of claim 10, wherein, According to the strategy of reordering the allocated memory space, the memory space for the first memory requirement comprises: Determining whether the sum of the available memory space of the memory can meet the first memory requirement; If so, the allocated memories are arranged in order from the start address or the end address of the memory by a memory moving module, and the allocated memory spaces are arranged continuously; After the arrangement, memory space is allocated for the first memory requirement from the remaining available memory space of the memory.

13. The memory management method of claim 10, wherein, According to the strategy of moving and releasing at least part of the allocated memory, the memory space for the first memory requirement comprises: Finding a second operator in the operator corresponding to the first memory requirement, wherein the second operator is the first operator whose state is undefined among the operators corresponding to the first memory requirement, and the undefined state is used to indicate that all memory requirements corresponding to the operator have not been allocated memory space; Finding and recording the memory requirements that have no corresponding relationship with the second operator from the memory requirements corresponding to the allocated memory space; select at least one part of the memory requirements from the recorded memory requirements, and move the stored data in the memory space corresponding to the selected at least one part of the memory requirements to another memory to release the memory space corresponding to the selected at least one part of the memory requirements, wherein the continuous available memory space formed by the memory space corresponding to the selected at least one part of the memory requirements and the available memory space in the memory is greater than or equal to the memory space required by the first memory requirement; allocate the memory space for the first memory requirement from the continuous available memory space.

14. The memory management method of claim 10, wherein, allocating the memory space for the first memory requirement by using the strategy of the splitting operator includes: finding a third operator in the operators corresponding to the first memory requirement and splitting the third operator, wherein the first memory requirement corresponds to other data of the third operator except input data; after the third operator is found and split, if the available memory space of the memory can meet the first memory requirement, reconstruct the sequence of the memory requirements and allocate the memory space for the first memory requirement.

15. The memory management method of claim 1, wherein, The network computing graph is a directed acyclic graph.

16. The memory management method of claim 1, wherein, The target memory block is a continuous memory space in the available memory space of the memory.

17. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the memory management method in any one of claims 1 to 16.

18. An electronic device, comprising: The electronic device includes: a memory storing a computer program; a processor connected in communication with the memory, and when the computer program is invoked, the memory management method in any one of claims 1 to 16 is executed.

Citation Information

Cited By

  • Program compiling method and device, electronic equipment, storage medium and program product

    CN121807276A