Memory allocation method and device based on reinforcement learning

CN116360962BActive Publication Date: 2026-09-25WUXI LINGXI BRAIN TECH CO LTD
View PDF 2 Cites -1 Cited by

Patent Information

Application Number
CN202111603884.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2026-09-25
Estimated Expiration
2041-12-24

AI Technical Summary

Benefits of technology

[0022]本公开实施例所提供的基于强化学习的内存分配方法、内存分配装置、电子设备及计算机可读介质的技术方案中,基于强化学习的价值迭代方式,对计算核对应的各算子进行计算核的内存分配,可以快速且容易实现对各计算核所对应的算子的内存分配,有效提高了内存分配效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116360962B_ABST
    Figure CN116360962B_ABST
Patent Text Reader

Abstract

The disclosure provides a memory allocation method based on reinforcement learning, which is applied to a many-core chip including a plurality of computing cores, each of which is configured with independent memory. The memory allocation method comprises: in a current round iteration, in a current memory allocation state of a current available memory, determining an allocable memory position corresponding to a current tensor according to tensor attribute information of the current tensor and context information of the current tensor, the context information including tensor attribute information of other tensors whose memory allocation time is later than that of the current tensor; performing memory allocation on the current tensor according to a state value score of a memory allocation state corresponding to each allocable memory position of the current tensor, and updating the current memory allocation state of the current available memory. The disclosure also provides a memory allocation device, an electronic device and a computer readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a memory allocation method and device based on reinforcement learning, an electronic device, and a computer-readable medium. Background Technology

[0002] Deep learning frameworks (such as TensorFlow or ONNX) typically use computation graphs to represent the computations of deep learning models (neural networks). For specific acceleration hardware, the neural network computation graph needs to be compiled by a compiler to generate an instruction stream that can run on the hardware. This hardware can be a many-core architecture chip with in-memory computing capabilities, which typically includes multiple compute cores (COREs).

[0003] During the compilation phase of a neural network computation graph, after the computation graph enters the compiler, tasks are allocated. Based on the computational complexity and memory access requirements of different operators in the computation graph, as well as the synchronization information between operators, different operators are assigned to different computational cores for execution. Each operator corresponds to at least one tensor.

[0004] Given the limited memory on the computational core, how to allocate the corresponding computational core memory to operator tensors in a reasonable and effective manner is a technical problem that urgently needs to be solved. Summary of the Invention

[0005] This disclosure provides a memory allocation method, memory allocation device, electronic device, and computer-readable medium based on reinforcement learning.

[0006] According to a first aspect of this disclosure, embodiments of this disclosure provide a memory allocation method based on reinforcement learning, which is applied to a many-core chip. The many-core chip includes multiple computing cores, each of which is configured with independent memory. The memory allocation method includes: in the current iteration, performing memory allocation operations sequentially on each operator corresponding to the current computing core, wherein each operator includes at least one tensor, and the tensor type of the tensor is the same as the operator type of the operator; the memory allocation operation includes: in the current memory allocation state of the currently available memory, determining the allocatable memory location corresponding to the current tensor based on the tensor attribute information and the context information of the current tensor; the current memory allocation state represents the memory allocation status of each tensor allocated in the current iteration, wherein the context information includes tensor attribute information of other tensors whose memory allocation time is after the current tensor, and the tensor attribute information includes tensor type, tensor size, tensor memory allocation time, and memory release time; and allocating memory to the current tensor based on the state value score of the current tensor in the memory allocation state corresponding to each allocatable memory location, and updating the current memory allocation state of the currently available memory.

[0007] In some embodiments, the step of allocating memory to the current tensor based on the state value score of the memory allocation state corresponding to each allocable memory location of the current tensor, and updating the current memory allocation state of the currently available memory, includes: for each allocable memory location corresponding to the current tensor, determining the state value score of the memory allocation state when selecting that allocable memory location to allocate memory to the current tensor; allocating memory to the current tensor based on the allocable memory location corresponding to the memory allocation state with the highest state value score, and updating the current memory allocation state of the currently available memory.

[0008] In some embodiments, determining the allocatable memory location corresponding to the current tensor based on the tensor attribute information and the context information of the current tensor includes: determining all free memory locations in the currently available memory under the current memory allocation state; removing free memory locations from all free memory locations that do not match the tensor type of the current tensor; and determining the allocatable memory location corresponding to the current tensor from the remaining free memory locations based on the tensor size, memory allocation time, and memory release time of the current tensor, as well as the tensor size, memory allocation time, and memory release time of other tensors in the context information.

[0009] In some embodiments, determining the allocatable memory location corresponding to the current tensor from the remaining free memory locations based on the current tensor's size, memory allocation time, memory release time, and the tensor sizes, memory allocation times, and memory release times of other tensors in the context information includes: performing memory allocation combinations on the current tensor and other tensors based on the current tensor's size, memory allocation time, memory release time, and the tensor sizes, memory allocation times, and memory release times of other tensors in the context information, with each memory allocation combination corresponding to an allocation situation of the current tensor and other tensors in the remaining free memory locations; in performing memory allocation combinations, when any two tensors meet the memory reuse condition, the two tensors are set adjacently, and the memory location of the tensor that releases memory first is set to be allocated to the other tensor after release; for each memory allocation combination, when the memory size required by the memory allocation combination is less than or equal to the total memory size corresponding to the remaining free memory locations, the free memory location corresponding to the current tensor in the memory allocation combination is determined as the allocatable memory location corresponding to the current tensor.

[0010] In some embodiments, before determining the allocatable memory location corresponding to the current tensor based on the tensor attribute information and the context information of the current tensor, the method further includes: under the current memory allocation state, checking whether the tensor of the allocated memory location meets the memory release condition; when the tensor of the allocated memory location meets the memory release condition, releasing the memory location corresponding to the tensor to the currently available memory; wherein, the memory release condition includes the memory release time of the tensor of the allocated memory location being equal to the memory allocation time of the current tensor.

[0011] In some embodiments, if memory allocation for the current tensor fails, the memory allocation method further includes: obtaining the state penalty value corresponding to each allocated tensor, wherein each allocated tensor includes the current tensor and tensors that have allocated memory locations before the memory location of the current tensor is allocated; and updating the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor.

[0012] In some embodiments, obtaining the state penalty value corresponding to each allocated tensor includes: obtaining the distance parameter between each allocated tensor and the current tensor; the distance parameter represents the distance between the allocated tensor and the current tensor in the memory allocation order; for each allocated tensor, determining the state penalty value corresponding to the allocated tensor based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality; wherein, the larger the distance parameter, the smaller the corresponding state penalty value.

[0013] In some embodiments, determining the state penalty value corresponding to the allocated tensor based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality includes: obtaining the state penalty value corresponding to the allocated tensor using a state penalty value function based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality; the state penalty value function includes: Sn = func(Dn, C); where Sn represents the state penalty value corresponding to the nth allocated tensor, Dn is the distance parameter between the nth allocated tensor and the current tensor, C is the penalty cardinality, func() represents the state penalty value function related to Dn and C, and func() is C raised to the power of Dn.

[0014] In some embodiments, updating the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor includes: updating the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor using a predetermined value function; wherein the predetermined value function includes: state(n)_value1 = state(n)_value0 - Sn; where state(n)_value0 represents the current state value score of the memory allocation state corresponding to the nth tensor, state(n)_value1 represents the updated state value score of the memory allocation state corresponding to the nth tensor, and Sn represents the state penalty value corresponding to the nth tensor.

[0015] In some embodiments, after updating the current memory allocation state, the memory allocation method further includes: recording the memory allocation information corresponding to the current tensor, wherein the memory allocation information includes the current memory allocation state and the corresponding state value score.

[0016] In some embodiments, if memory allocation for the current tensor fails, the memory allocation method further includes: initializing the memory allocation state of the available memory of the current computing core; and performing the next iteration based on the memory allocation information corresponding to each tensor recorded in advance, wherein the memory allocation information corresponding to the tensor includes the memory allocation state corresponding to the tensor and the corresponding state value score.

[0017] In some embodiments, determining the state value score of the memory allocation state when selecting the allocatable memory location to allocate memory to the current tensor includes: when the current iteration is the first iteration, determining the state value score of the memory allocation state when selecting the allocatable memory location to allocate memory to the current tensor based on the initialized state value score; when the current iteration is not the first iteration, determining the state value score of the memory allocation state when selecting the allocatable memory location to allocate memory to the current tensor based on the memory allocation information corresponding to each tensor recorded in historical iterations; wherein, the memory allocation information corresponding to the tensor includes the memory allocation state corresponding to the tensor and the corresponding state value score.

[0018] In some embodiments, determining the state value score of the memory allocation state when selecting the allocatable memory location to allocate memory to the current tensor based on the memory allocation information corresponding to each tensor recorded in historical iterations includes: when there is a memory allocation state in the recorded memory allocation information that is the same as the memory allocation state corresponding to the current tensor at the allocatable memory location, determining the state value score of the memory allocation state in the recorded memory allocation information; when there is no memory allocation state in the recorded memory allocation information that is the same as the memory allocation state corresponding to the current tensor at the allocatable memory location, determining the state value score of the memory allocation state corresponding to the current tensor at the allocatable memory location based on the initialized state value score.

[0019] According to a second aspect of this disclosure, embodiments of this disclosure provide a memory allocation device based on reinforcement learning. This memory allocation device is applied to a many-core chip, which includes multiple computing cores, each of which is configured with independent memory. The memory allocation device includes: an allocation module configured to sequentially perform memory allocation operations on each operator corresponding to the current computing core in the current iteration. Each operator includes at least one tensor, and the tensor type of the tensor is the same as the operator type of the operator. The allocation module includes: a position determination submodule, used to determine the position based on the tensor attribute information of the current tensor, given the current memory allocation state of the currently available memory. The context information of the current tensor determines the allocatable memory location corresponding to the current tensor in the currently available memory. The current memory allocation status represents the memory allocation status of each tensor allocated in the current iteration. The context information includes tensor attribute information of other tensors whose memory allocation time is after the current tensor. The tensor attribute information includes tensor type, tensor size, tensor memory allocation time, and memory release time. The memory allocation submodule is used to allocate memory to the current tensor according to the state value score of the current tensor's memory allocation status at each allocatable memory location, and update the current memory allocation status of the currently available memory.

[0020] According to a third aspect of this disclosure, an embodiment of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the memory allocation method provided in any of the above embodiments.

[0021] According to a fourth aspect of this disclosure, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the memory allocation method provided in any of the above embodiments.

[0022] The technical solutions of the memory allocation method, memory allocation device, electronic device and computer-readable medium based on reinforcement learning provided in the embodiments of this disclosure, based on the value iteration method of reinforcement learning, allocate memory for each operator corresponding to the computational core, which can quickly and easily realize the memory allocation for the operators corresponding to each computational core, and effectively improve the memory allocation efficiency.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0024] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0025] Figure 1 A flowchart illustrating a memory allocation method based on reinforcement learning provided in an embodiment of this disclosure;

[0026] Figure 2 for Figure 1 A flowchart illustrating a specific implementation of step S11;

[0027] Figure 3 for Figure 2 A flowchart illustrating a specific implementation of step S113;

[0028] Figure 4 for Figure 1 A flowchart illustrating a specific implementation of step S12;

[0029] Figure 5 A flowchart illustrating another memory allocation method provided in this embodiment of the disclosure;

[0030] Figure 6 A flowchart illustrating another memory allocation method provided in this embodiment of the disclosure;

[0031] Figure 7 A flowchart illustrating another memory allocation method provided in this embodiment of the disclosure;

[0032] Figure 8 A flowchart illustrating another memory allocation method provided in this embodiment of the disclosure;

[0033] Figure 9 This is a schematic diagram of the structure of a memory allocation device provided in an embodiment of the present disclosure;

[0034] Figure 10 This is a schematic diagram of another memory allocation device provided in an embodiment of the present disclosure;

[0035] Figure 11 This is a schematic diagram of another memory allocation device provided in an embodiment of the present disclosure;

[0036] Figure 12 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0037] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0038] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0039] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0040] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0041] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0042] This disclosure provides a memory allocation method based on reinforcement learning. The memory allocation method is implemented using a memory allocation device, which can be implemented by software and / or hardware. The memory allocation device can be integrated into an electronic device. For example, the type of electronic device can be a laptop, computer, server, mobile phone, tablet computer (PAD), etc., and is not particularly limited in this disclosure.

[0043] In this embodiment of the disclosure, the memory allocation method is applied to a many-core chip, which may include multiple computing cores (COREs), each of which is configured with independent memory.

[0044] In this embodiment, the memory allocation method is used for memory allocation in a neural network computation graph, specifically for allocating memory for the tensors of operators allocated to the computation kernel. The neural network computation graph can include multiple operators, which are the basic computational units constituting the neural network. These operators are represented by nodes in the neural network computation graph. The types of operators can be, for example, convolution, pooling, etc. An operator can include one or more tensors, which can be the operator's input tensor, output tensor, etc.

[0045] For many-core chips, each computing core can have multiple execution engines of different types running simultaneously. Different operators are suitable for running on different execution engines; that is, operators of different types are suitable for running on different execution engines. For example, some execution engines have special acceleration operations for convolution, while others are suitable for vector operations. Once the memory of an operator executed on a certain type of execution engine is released, it can be reused by other operators of the same type. Tensors have the same tensor type as the operator type of their respective operators, and therefore, the same execution engine.

[0046] Figure 1 A flowchart illustrating a reinforcement learning-based memory allocation method provided in this disclosure is shown below. Figure 1This disclosure provides a memory allocation method based on reinforcement learning, including: in the current iteration, performing memory allocation operations on each operator corresponding to the current computation kernel in sequence, each operator including at least one tensor, the tensor type of the tensor being the same as the operator type of the operator, and the memory allocation operation including the following steps: steps S11 to S12.

[0047] In this embodiment of the disclosure, the current computing core is any one of the many-core chips that needs to allocate memory for the corresponding operator tensor. Each operator corresponding to the current computing core is an operator allocated to the current computing core. Each operator includes at least one tensor, and each tensor has corresponding tensor attribute information. The tensor attribute information represents information used to characterize the tensor's feature attributes. The tensor attribute information may include tensor type, tensor size, memory allocation time, and memory release time. The tensor type is the same as the type of the corresponding operator. The tensor size is the amount of memory required by the tensor. The tensor's memory allocation time (phase) refers to the allocation time for allocating the memory corresponding to the tensor, and the memory release time (freephase) refers to the release time for releasing the memory corresponding to the tensor.

[0048] In this embodiment of the disclosure, the memory allocation order of each tensor can be determined according to the order of memory allocation time of each tensor, and memory allocation operations can be performed on each tensor in sequence according to the memory allocation order of each tensor.

[0049] Step S11: Under the current memory allocation state of the currently available memory, determine the allocatable memory location corresponding to the current tensor in the currently available memory based on the tensor attribute information and the context information of the current tensor.

[0050] In this embodiment of the disclosure, the current tensor refers to the tensor that currently requires memory allocation according to the memory allocation order. In step S11, for the current tensor in the memory allocation order, under the current memory allocation state of the currently available memory, the allocatable memory location corresponding to the current tensor in the currently available memory can be determined based on the current available memory situation under the current memory allocation state, as well as the tensor attribute information and the context information of the current tensor. The allocatable memory location refers to the free memory location that can be allocated to the current tensor under the current memory allocation state.

[0051] In this embodiment of the disclosure, the context information includes tensor attribute information of other tensors whose memory allocation time follows the current tensor. These other tensors whose memory allocation time follows the current tensor can be m (m greater than or equal to 1) tensors located after the current tensor in the memory allocation order. For example, m can be set to 1, 2, 3, 4, or 5. The specific value of m can be set according to actual conditions and needs, and this embodiment of the disclosure does not impose any special restrictions on it.

[0052] In this embodiment of the disclosure, the current memory allocation state represents the memory allocation status of each tensor allocated in the current iteration. In some embodiments, the current memory allocation state may include a mapping relationship between allocated tensors, corresponding allocated memory locations, and information on available memory locations at the time of allocation. This mapping relationship can be used to represent the memory occupancy of tensors allocated in the current iteration within the available memory of the current computing core. Here, an allocated tensor refers to a tensor whose memory was successfully allocated before the current tensor's memory was allocated in the memory allocation order; the corresponding allocated memory location refers to the memory location allocated to that tensor; and the information on available memory locations at the time of allocation refers to all available memory locations corresponding to that tensor at the time of memory allocation.

[0053] For example, suppose the tensors corresponding to the current computing core include tensor A, tensor B, and tensor C, and the memory allocation order is from tensor A to tensor B, and then to tensor C. The current tensor is tensor B, and the allocated tensor is tensor A. The available memory locations of the current computing core include memory location 1, memory location 2, memory location 3, memory location 4, and memory location 5. The memory locations allocated to tensor A include memory locations 1 and 2. Then the current memory allocation state includes tensor A, memory locations 1 and 2, and the mapping relationship of the allocatable memory locations 1-5 corresponding to tensor A during allocation. The allocatable memory locations corresponding to the current tensor B include memory locations 3-5.

[0054] In some embodiments, the above mapping relationship can be represented by encoding the identifier of the allocated tensor, the corresponding allocated memory location, and the information of the allocatable memory location at the time of allocation, with each mapping relationship corresponding to an allocated tensor.

[0055] In some embodiments, when determining the allocatable memory location corresponding to the current tensor in the currently available memory, a resource preemption principle is followed. For example, the resource preemption principle includes: assuming that all tensors are of two types, A and B, for memory locations allocated of type A, after being freed, they can only be allocated to tensors of type A. When allocating memory for a tensor of type A, a memory location is searched for and allocated from the memory locations corresponding to the freed tensors of type A and memory locations that have not been used by any tensors. Similarly, for memory locations allocated of type B, after being freed, they can only be allocated to tensors of type B. When allocating memory for a tensor of type B, a memory location is searched for and allocated from the memory locations corresponding to the freed tensors of type B and memory locations that have not been used by any tensors.

[0056] Figure 2 for Figure 1A flowchart illustrating a specific implementation of step S11 is shown in some embodiments, such as... Figure 2 As shown, step S11 may further include: steps S111 to S113.

[0057] Step S111: Determine all free memory locations in the currently available memory under the current memory allocation state.

[0058] Free memory locations refer to memory locations that are currently unoccupied and have not been allocated.

[0059] Step S112: Remove free memory locations from all free memory locations that do not match the tensor type of the current tensor.

[0060] In this context, a free memory location that does not match the tensor type of the current tensor refers to a tensor whose tensor type was previously used at that location, and this free memory location is created by releasing memory locations of tensors of other tensor types. These other tensor types are those different from the current tensor's tensor type. For example, if the current tensor's tensor type is A, and the free memory location X was previously allocated to a tensor of tensor type B, then the free memory location X does not match the current tensor's tensor type A.

[0061] Step S113: Based on the current tensor's size, memory allocation time, and memory release time, as well as the tensor sizes, memory allocation times, and memory release times of other tensors in the context information, determine the allocatable memory location corresponding to the current tensor from the remaining free memory locations.

[0062] Figure 3 for Figure 2 A flowchart illustrating a specific implementation of step S113 is shown in some embodiments, such as... Figure 3 As shown, step S113 may further include: steps S113a to S113d.

[0063] Step S113a: Based on the current tensor's size, memory allocation time, and memory release time, as well as the size, memory allocation time, and memory release time of other tensors in the context information, perform memory allocation combination on the current tensor and other tensors.

[0064] Each memory allocation combination corresponds to an allocation of the current tensor and other tensors in the remaining free memory locations. The allocation can include the free memory locations pre-allocated to each tensor in the remaining free memory locations, and the storage order of each tensor in the remaining free memory locations.

[0065] In step S113b, during memory allocation and combination, when any two tensors meet the memory reuse condition, the two tensors are set to be adjacent, and the memory location of the tensor that releases memory first is set to be allocated to the other tensor after release.

[0066] Among them, the memory reuse condition includes the memory release time of one tensor being earlier than (less than or equal to) the memory allocation time of another tensor.

[0067] Step S113c: For each memory allocation combination, determine the memory size required for that memory allocation combination.

[0068] When there are no tensors in a memory allocation combination that satisfy the memory reuse condition, the memory size required for the memory allocation combination is determined by the sum of the memory sizes required by each tensor (tensor size). When there are multiple tensors in a memory allocation combination that satisfy the memory reuse condition, the memory size required by these multiple tensors is determined by the memory size required by the tensor with the largest tensor size among them. Furthermore, the memory size required by the tensor with the largest tensor size among these multiple tensors is the total memory size required by these multiple tensors.

[0069] Step S113d: When the memory size required by the memory allocation combination is less than or equal to the total memory size corresponding to the remaining free memory locations, the free memory location corresponding to the current tensor in the memory allocation combination is determined as the allocatable memory location corresponding to the current tensor.

[0070] When performing memory allocation and combination, the principle of reusing memory as much as possible can be followed. Memory reuse means that when the memory release time of one tensor is earlier than (less than or equal to) the memory allocation time of another tensor, the two tensors are set adjacently during combination, and multiple adjacent and contiguous memory regions are allocated to the two tensors during allocation. Furthermore, the memory region corresponding to one tensor can be reused by the other tensor. The memory region of the tensor that releases memory first is set to be allocated to the other tensor after release. The size of the memory allocated to the two tensors is the size required by the tensor with the larger memory. This effectively improves memory utilization and minimizes memory fragmentation.

[0071] For example, in the available memory of size 5 memory units, memory corresponding to tensor A, tensor B and tensor C are allocated in the memory allocation order respectively. The size of tensor A is 2 memory units, the size of tensor B is 1 memory unit, and the size of tensor C is 4 memory units. Tensor A is the current tensor. The available memory locations include memory location 1, memory location 2, memory location 3, memory location 4 and memory location 5, and the memory size of each memory location is 1 memory unit.

[0072] When memory is allocated in memory locations 1-5 in the order of tensor A to tensor B to tensor C, tensor A will be allocated to memory locations 1 and 2, tensor B will be allocated to memory location 3, and the remaining memory locations 4 and 5 will not be able to meet the memory requirement of tensor C, which is 4 memory units in size. In other words, tensor C will not be allocated successfully.

[0073] When memory is allocated in memory locations 1-5 in the order of tensor B to tensor A to tensor C, tensor A, which has a size of 2 memory units, will be allocated to memory locations 2 and 3, and tensor B, which has a size of 1 memory unit, will be allocated to memory location 1. The memory release time of tensor A is less than or equal to the memory allocation time of tensor C. Therefore, when allocating memory for tensor C, the memory allocated to tensor A has been released, that is, the memory at memory locations 2 and 3 is released. At this time, memory locations 2-5 can be allocated to tensor C, thereby achieving memory reuse. It can be understood that the released memory locations 2 and 3 are reused by tensor C, and the memory size allocated to the combination of tensors A and C is equal to the memory size required by tensor C.

[0074] When memory is allocated in memory locations 1-5 in the order of tensor A to tensor C to tensor B, tensor A will be allocated to memory locations 1 and 2, and tensor B will be allocated to memory location 5. The memory release time of tensor A is less than or equal to the memory allocation time of tensor C. Therefore, when the memory of tensor C is allocated, the memory allocated to tensor A has been released, that is, the memory of memory locations 1 and 2 is released. At this time, memory locations 1-4 can be allocated to tensor C for use, thereby realizing memory reuse.

[0075] For other combinations of tensor A, tensor B, and tensor C in memory locations 1-5, refer to the descriptions of the cases listed above; they will not be repeated here. Based on the allocation under different combinations, the allocatable memory location of tensor A in the current memory state can be determined.

[0076] In this embodiment of the disclosure, when the current iteration is the first iteration, before performing memory allocation operations on each operator corresponding to the current computing core in the current iteration, the memory allocation method further includes: setting reinforcement learning target conditions and initializing the memory allocation state of the available memory of the current computing core.

[0077] The reinforcement learning objective conditions include at least one of the following: completing the memory allocation for the last tensor in the memory allocation sequence, or reaching a predetermined maximum number of iterations (e.g., 100 iterations). Before the first iteration, in the initialized memory allocation state, all available memory locations are in an unallocated state.

[0078] In this embodiment of the disclosure, the available memory of the current computing core may include multiple discrete memory regions (locations) or multiple consecutive memory regions (locations), which can be determined according to the actual situation.

[0079] In some embodiments, before determining the allocatable memory location corresponding to the current tensor based on the tensor attribute information and the context information of the current tensor, i.e. before step S11, the memory allocation method further includes: in the current memory allocation state, checking whether the tensor of the allocated memory location meets the memory release condition. The memory release condition includes that the memory release time of the tensor of the allocated memory location is equal to the memory allocation time of the current tensor.

[0080] When a tensor with allocated memory locations meets the memory release condition (i.e., the memory release time of the tensor with allocated memory locations is equal to the memory allocation time of the current tensor), the memory location corresponding to that tensor is released to the currently available memory. When a tensor with allocated memory locations does not meet the memory release condition (i.e., the memory release time of the tensor with allocated memory locations is not equal to the memory allocation time of the current tensor), the system continues to check whether the tensor with the next allocated memory location meets the memory release condition, until all tensors with allocated memory locations have been traversed.

[0081] Step S12: Allocate memory for the current tensor based on the state value score of the memory allocation state corresponding to each allocatable memory location, and update the current memory allocation state of the currently available memory.

[0082] Figure 4 for Figure 1 A flowchart illustrating a specific implementation of step S12 is shown in some embodiments, such as... Figure 2 As shown, step S12 may further include steps S121 and S122.

[0083] Step S121: For each allocatable memory location corresponding to the current tensor, determine the state value score of the memory allocation state when selecting that allocatable memory location to allocate memory to the current tensor.

[0084] The memory allocation state of a tensor differs depending on its allocation to different allocable memory locations. The memory allocation state represents the memory allocation status of each allocated tensor. The state value score of the current tensor when allocated to any allocable memory location represents the value of its memory allocation state at that location. Initially, the state value score for each memory allocation state can be initialized to a predetermined value (e.g., 0). For each tensor, the state value score of its corresponding memory allocation state can be iteratively updated if memory allocation fails for that tensor or for subsequent tensors. The update method can be decremental, and the decrement value each time can be determined based on the distance between the tensor and the tensors that failed to allocate memory in the memory allocation order. Specifically, the distance between two adjacent tensors in the memory allocation order can be set to a predetermined distance, thus determining the distance between any two tensors in the memory allocation order. The larger the distance, the smaller the decrement value in each update.

[0085] In the current memory allocation state, each allocatable memory location corresponding to the current tensor corresponds to a possible allocation action. For each possible allocation action, obtain the state value score of the next memory allocation state reached after executing the allocation action. Executing the allocation action means selecting the allocatable memory location corresponding to the allocation action to allocate memory to the current tensor, thereby obtaining a new memory allocation state. Each allocation action corresponds to a new memory allocation state.

[0086] For example, suppose the tensors corresponding to the current computing core include tensor A, tensor B, and tensor C, and the memory allocation order is from tensor A to tensor B, and then to tensor C. The current tensor is tensor B, and the allocated tensor is tensor A. The available memory locations of the current computing core include memory location 1, memory location 2, memory location 3, memory location 4, and memory location 5. Then the current memory allocation state includes tensor A, memory locations 1 and 2, and the mapping relationship of the allocatable memory locations 1-5 corresponding to tensor A during allocation. The allocatable memory locations corresponding to the current tensor B include memory location 3, memory location 4, and memory location 5.

[0087] The memory allocation state reached after performing the action of allocating memory to the current tensor B at memory location 3 includes: tensor A, memory locations 1 and 2, and the mapping relationship of allocable memory locations 1-5 corresponding to tensor A at the time of allocation; and tensor B, memory location 3, and the mapping relationship of allocable memory locations 3-5 corresponding to tensor B at the time of allocation. Similarly, the memory allocation state reached after performing the action of allocating memory to the current tensor B at memory location 4 includes: tensor A, memory locations 1 and 2, and the mapping relationship of allocable memory locations 1-5 corresponding to tensor A at the time of allocation; and tensor B, memory location 4, and the mapping relationship of allocable memory locations 3-5 corresponding to tensor B at the time of allocation. The memory allocation state reached after performing the action of allocating memory to the current tensor B at memory location 5 includes: tensor A, memory locations 1 and 2, and the mapping relationship of allocable memory locations 1-5 corresponding to tensor A at the time of allocation; and tensor B, memory location 5, and the mapping relationship of allocable memory locations 3-5 corresponding to tensor B at the time of allocation.

[0088] In some embodiments, when the current iteration is the first iteration (first iteration), the memory allocation state generated by each new allocation action is a completely new (never before generated in the past) memory allocation state. In this case, the state value score of the memory allocation state when allocating memory to the current tensor at the selected allocable memory location is determined based on the initialized state value score. Further, the state value score of the memory allocation state when allocating memory to the current tensor at the selected allocable memory location is the initialized state value score. The specific value of the initialized state value score can be set according to actual needs; for example, the default initialized state value score is 0.

[0089] In some embodiments, when the current iteration is not the first iteration (e.g., the second iteration, the third iteration, etc.), the memory allocation state resulting from an allocation action may have appeared in previous iterations. In this case, based on the memory allocation information corresponding to each allocated tensor recorded in previous iterations, the state value score of the memory allocation state when selecting the allocatable memory location to allocate memory to the current tensor is determined. The memory allocation information corresponding to the allocated tensor includes the memory allocation state corresponding to the allocated tensor and the corresponding state value score.

[0090] In this embodiment, when there is a memory allocation state in the memory allocation information recorded in the historical iterations that is the same as the memory allocation state corresponding to the current tensor at the allocable memory location, the state value score corresponding to the memory allocation state in the memory allocation information is determined as the state value score of the memory allocation state corresponding to the current tensor at the allocable memory location.

[0091] In this embodiment, when there is no memory allocation state in the memory allocation information recorded in the historical iterations that is the same as the memory allocation state corresponding to the current tensor at the allocable memory location, the state value score of the memory allocation state corresponding to the current tensor at the allocable memory location is determined according to the initialized state value score. Further, the state value score of the memory allocation state corresponding to the current tensor at the allocable memory location is the initialized state value score.

[0092] It should be noted that the same memory allocation state means that the allocated tensors corresponding to the two memory allocation states are the same (with the same identifier), the allocated memory locations corresponding to the same tensors are the same, and the allocatable memory location information corresponding to the same tensors is the same during allocation.

[0093] It is understandable that a historical iteration refers to any iteration performed before the current iteration. For example, if the current iteration is the fourth iteration, then the historical iteration could be the third iteration, the second iteration, or the first iteration.

[0094] Step S122: Allocate memory for the current tensor based on the allocable memory location corresponding to the memory allocation state with the highest state value score, and update the current memory allocation state of the currently available memory.

[0095] In this embodiment of the disclosure, after determining the state value score of the memory allocation state corresponding to each allocatable memory location of the current tensor, the memory allocation state with the highest state value score is selected as the target memory allocation state, and the allocatable memory location corresponding to the target memory allocation state is used as the memory location currently allocated to the current tensor, so as to allocate memory to the current tensor.

[0096] In some embodiments, when there are multiple memory allocation states with the highest state value scores, any one of the memory allocation states can be randomly selected as the target memory allocation state.

[0097] In this embodiment of the disclosure, when allocating memory to the current tensor by selecting the allocable memory location corresponding to the memory allocation state with the highest state value score, the current memory allocation state of the currently available memory is updated.

[0098] For example, suppose the tensors corresponding to the current computing core include tensor A, tensor B, and tensor C, the current tensor is tensor B, the allocated tensors include tensor A, and the available memory locations of the current computing core include memory location 1, memory location 2, memory location 3, memory location 4, and memory location 5. Before allocating memory for the current tensor B, the current memory allocation state includes tensor A, memory locations 1 and 2, and the mapping relationship of allocable memory locations 1-5 corresponding to tensor A at the time of allocation. When allocating memory for the current tensor B, the allocable memory locations include memory locations 3-5, among which the state value score corresponding to allocable memory location 3 is the highest. Therefore, memory location 3 is allocated to the current tensor B. At this time, the current memory allocation state of the currently available memory is updated. After the update, the current memory allocation state includes the allocated tensor A, memory locations 1 and 2, and the mapping relationship of allocable memory locations 1-5 corresponding to tensor A at the time of allocation, as well as the allocated tensor B, memory location 3, and the mapping relationship of allocable memory locations 3-5 corresponding to tensor B at the time of allocation.

[0099] In this embodiment of the disclosure, when allocating memory to the current tensor based on the allocatable memory location corresponding to the memory allocation state with the highest state value score, if the memory size corresponding to the allocatable memory location corresponding to the memory allocation state with the highest state value score does not meet the memory requirements of the current tensor, then the memory allocation to the current tensor fails; if the memory size corresponding to the allocatable memory location corresponding to the memory allocation state with the highest state value score meets the memory requirements of the current tensor, then the memory allocation to the current tensor succeeds.

[0100] For example, suppose the current tensor C requires 4 memory units, while the memory allocated to memory location 4-5 of tensor C is 2 memory units, which is less than the current tensor C requires. This means that the memory allocated to memory location 4-5 of the current tensor C does not meet the current tensor's memory requirements.

[0101] The memory allocation method based on reinforcement learning provided in this embodiment allocates memory for each operator corresponding to a computational core based on the value iteration method of reinforcement learning. This method can quickly and easily allocate memory for each operator corresponding to a computational core, effectively improving memory allocation efficiency.

[0102] Figure 5 This is a flowchart illustrating another memory allocation method provided in embodiments of the present disclosure. In some embodiments, such as... Figure 5 As shown, in the event that memory allocation for the current tensor fails, the memory allocation method further includes steps S131 to S132.

[0103] Step S131: Obtain the state penalty value corresponding to each of the allocated tensors.

[0104] It is understood that the allocated tensors include the current tensor, as well as the tensors whose memory locations were allocated before the current tensor was allocated.

[0105] The state penalty value represents the magnitude of the impact on the state value of the memory allocation state of the current tensor and the previously allocated tensors when memory allocation of the current tensor fails. The further the allocated tensor is from the current tensor that failed to be allocated in the memory allocation order, the smaller the corresponding state penalty value.

[0106] In some embodiments, obtaining the state penalty value corresponding to each allocated tensor includes: obtaining the distance parameters between each allocated tensor and the current tensor, where the distance parameters represent the distance between the allocated tensors and the current tensor in the memory allocation order; for each allocated tensor, determining the state penalty value corresponding to the allocated tensor based on the distance parameters between the allocated tensor and the current tensor and the penalty cardinality. Wherein, the larger the distance parameter, the smaller the corresponding state penalty value.

[0107] In some embodiments, determining the state penalty value corresponding to the allocated tensor based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality includes: obtaining the state penalty value corresponding to the allocated tensor using a state penalty value function based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality; wherein, the state penalty value function includes: Sn = func(Dn, C), Sn represents the state penalty value corresponding to the nth allocated tensor, Dn is the distance parameter between the nth allocated tensor and the current tensor, C is the penalty cardinality, the penalty cardinality C is a preset constant, and the specific value of the penalty cardinality C can be set as needed, and func() represents the state penalty value function related to Dn and C. In some embodiments, func() can be C raised to the power of Dn.

[0108] Step S132: Update the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor.

[0109] In some embodiments, updating the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor includes: updating the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor using a predetermined value function.

[0110] The predetermined value function includes: state(n)_value1 = state(n)_value0 - Sn. Here, state(n)_value0 represents the current state value score of the memory allocation state corresponding to the nth tensor, which is the state value score of the memory allocation state determined in step S121. state(n)_value1 represents the updated state value score of the memory allocation state corresponding to the nth tensor, and Sn represents the state penalty value corresponding to the nth tensor; that is, the updated state value score is the difference between the state value score before the update and the state penalty value.

[0111] For example, the penalty base C is set to 0.1. In the current iteration, the current tensor is the 10th tensor in the memory allocation order. The distance parameter between two adjacent tensors in the memory allocation order is defined as 1 unit distance. The distance parameter between the current tensor and the 10th tensor is D10 = 0. When memory allocation for the 10th tensor fails, the state value score of the memory allocation state corresponding to the 10th tensor is updated to state(10)_value1 = state(10)_value0 - (0.1). 0 =state(10)_value0-1,(0.1) 0 That is, the state penalty value corresponding to the current tensor (the 10th tensor). state(10)_value0 can be determined according to the state value score determined in step S121. In other words, the current state value score of the memory allocation state corresponding to the 10th tensor is the state value score of the memory allocation state determined in step S121.

[0112] Similarly, the state value score of the memory allocation state corresponding to the 9th tensor allocated in the current iteration is updated to state(9)_value 1 = state(9)_value 0 - (0.1). 1 (0.1) 1 This is the state penalty value corresponding to the 9th tensor. The state value score of the memory allocation state corresponding to the 8th tensor is updated to state(8)_value 1 = state(8)_value 0 - (0.1). 2 (0.1) 2 This is the state penalty value corresponding to the 8th tensor. The state value score of the memory allocation state corresponding to the 7th tensor is updated to state(7)_value 1 = state(7)_value 0 - (0.1). 3 (0.1) 3 This is the state penalty value corresponding to the 7th tensor, and so on.

[0113] According to the value function above, when the memory allocation of the current tensor fails, the further away the other allocated tensors are from the current tensor that failed in the memory allocation order, the less the state value score of the memory allocation state corresponding to the other allocated tensors will be affected, that is, the smaller the state penalty value will be.

[0114] It should be noted that the func() function can be predefined according to actual needs, as long as it can satisfy the condition that the further away from the current tensor where the allocation failed, the less the impact on the state value score of the corresponding memory allocation state is, that is, the smaller the state penalty value.

[0115] Understandably, whether to update the state value score of the memory allocation state corresponding to each allocated tensor (including the current tensor) depends on whether the memory allocation of the current tensor is successful. If the memory allocation of the current tensor is successful, the state value score of the memory allocation state corresponding to each allocated tensor (including the current tensor) remains unchanged, where the state value score of the current memory allocation state corresponding to the current tensor is the same as the state value score of the memory allocation state determined in step S121. If the memory allocation of the current tensor fails, the state value score of the memory allocation state corresponding to each allocated tensor (including the current tensor) is updated according to the value function.

[0116] Furthermore, in the value function, the current state value score state(n)_value0 of the memory allocation state is the state value score of the memory allocation state determined in step S121. The state value score of the memory allocation state determined in step S121 depends on whether there is a memory allocation state with the same memory allocation state in the historical iteration.

[0117] For example, if the current memory allocation state of the current tensor in the current iteration is the same as the memory allocation state of the current tensor in the previous iteration, then the current state value score of the current memory allocation state, state(n)_value0, is the same as the state value score of the memory allocation state of the current tensor in the previous iteration. If memory allocation for the current tensor fails, the state value score of the current memory allocation state is updated to state(n)_value0-1. For example, if the current iteration is the second iteration, and there was a memory allocation state with the same state as the current one in the first iteration, and the state value score of that memory allocation state was -0.5 in the first iteration, and memory allocation fails in that state in the second iteration, then the state value score of that memory allocation state is updated to -0.5-1 = -1.5.

[0118] In other words, in multiple iterations, the state value score of each generated memory allocation state may be updated once or multiple times during the iteration.

[0119] In some embodiments, a memory allocation state record table can be pre-set to record memory allocation states and their corresponding state value scores generated during the iteration process. Initially, this record table is empty; during iteration, it is recorded each time a new memory allocation state is generated, and the record table is updated with the memory allocation state and its corresponding state value score based on the memory allocation situation during multiple iterations. For a newly generated memory allocation state, if the record table does not record the same memory allocation state, the state value score of the newly generated memory allocation state is the initialized state value score, which can be set to 0 by default.

[0120] Figure 6 A flowchart illustrating another memory allocation method provided in this disclosure embodiment is shown below. Figure 6 As shown, in some embodiments, after updating the current memory allocation state, the memory allocation method further includes step S14.

[0121] Step S14: Record the memory allocation information corresponding to the current tensor. The memory allocation information includes the current memory allocation state corresponding to the current tensor and the corresponding state value score.

[0122] In some embodiments, if memory allocation of the current tensor is successful, the recorded state value score corresponding to the current memory allocation state is the same as the state value score corresponding to the memory allocation state determined in step S121. If memory allocation of the current tensor fails, the recorded state value score corresponding to the current memory allocation state is the updated state value score of the memory allocation state in steps S131-S132.

[0123] In some embodiments, in each iteration, memory allocation information of each allocated tensor in the current iteration is recorded so that when memory allocation of a tensor fails in the current iteration, the next iteration is entered based on the memory allocation information of each tensor recorded in the current iteration and previous iterations, and the state value score of each memory allocation state is iterated.

[0124] Figure 7 This is a flowchart illustrating a memory allocation method provided in an embodiment of the present disclosure, as shown below. Figure 7 As shown, in some embodiments, if memory allocation for the current tensor fails, the memory allocation method further includes steps S15 and S16.

[0125] Step S15: Initialize the memory allocation state of the available memory of the current computing core.

[0126] In step S15, the memory allocation state of the available memory of the current computing core is initialized, that is, the mapping relationship between each tensor and the allocated memory location in this iteration is cleared, while the memory allocation information of each tensor recorded in this iteration is retained, that is, the memory allocation state and the corresponding state value score.

[0127] Step S16: Perform the next iteration based on the memory allocation information corresponding to each tensor that has been recorded in advance.

[0128] The memory allocation information corresponding to the tensor includes the memory allocation state of the tensor and the corresponding state value score.

[0129] In step S16, the next iteration is performed based on the memory allocation information of each tensor recorded in the current iteration and the memory allocation information of each tensor recorded in previous iterations. In the next iteration, the above memory allocation method is repeated; details will not be elaborated here.

[0130] Figure 8 This is a flowchart illustrating a memory allocation method provided in an embodiment of the present disclosure, as shown below. Figure 8 As shown, in some embodiments, after updating the current memory allocation state of the currently available memory, i.e. after step S12, the memory allocation method further includes step S17.

[0131] Step S17: Determine whether the reinforcement learning objective conditions are met.

[0132] In the current iteration, when the memory allocation of the last tensor in the memory allocation sequence is successfully completed, it indicates that all tensors corresponding to the current computation kernel have been successfully allocated memory in the current iteration. Therefore, the reinforcement learning objective condition is deemed met, allocation completion information is returned, and the process ends. The allocation completion information includes at least the mapping relationship between each tensor and its allocated memory location.

[0133] In the current iteration, if the current tensor is not the last tensor in the memory allocation order, it is determined that the reinforcement learning objective condition is not met. If the memory allocation of the current tensor is successful, the memory allocation of the next tensor continues. If the memory allocation of the current tensor fails, and the current iteration is not the last iteration, i.e. the number of iterations has not reached the maximum number of iterations, then after updating the state value score using steps S131-S132 and recording the memory allocation state and state value score of the allocated tensors using step S14, steps S15-S16 are performed to enter the next iteration.

[0134] In the current iteration, if the current tensor is the last tensor in the memory allocation order, but memory allocation for the current tensor fails, and the current iteration is not the last iteration (i.e., the number of iterations has not reached the maximum number of iterations), then it is determined that the reinforcement learning objective condition is not met. After updating the state value score using steps S131-S132 and recording the memory allocation state and state value score of the allocated tensor using step S14, steps S15-S16 are performed to enter the next iteration.

[0135] In the current iteration, if memory allocation for the current tensor fails, and the current iteration is the last iteration (i.e., the maximum number of iterations has been reached), then it is determined that the reinforcement learning objective condition is met, the allocation failure information is returned, and the process ends.

[0136] Figure 9 This is a schematic diagram of a memory allocation device provided in an embodiment of the present disclosure, with reference to... Figure 9 The memory allocation device 400 is used to implement the above-described reinforcement learning-based memory allocation method. The memory allocation device 400 may include an allocation module configured to perform memory allocation operations sequentially on each operator corresponding to the current computational kernel in the current iteration. Each operator includes at least one tensor, and the tensor type is the same as the operator type of the operator it belongs to. The allocation module includes a position determination submodule 401 and a memory allocation submodule 403.

[0137] The location determination submodule 401 is used to determine the allocatable memory location corresponding to the current tensor in the current available memory based on the tensor attribute information and the context information of the current tensor, under the current memory allocation state of the current available memory. The current memory allocation state represents the memory allocation status of each tensor that has been allocated in the current iteration. The context information includes the tensor attribute information of other tensors whose memory allocation time is after the current tensor. The tensor attribute information includes tensor type, tensor size, tensor memory allocation time and memory release time.

[0138] The memory allocation submodule 403 is used to allocate memory to the current tensor based on the state value score of the memory allocation state corresponding to each allocatable memory location of the current tensor, and update the current memory allocation state of the currently available memory.

[0139] In some embodiments, the allocation module further includes a value determination submodule 402, which is used to determine the state value score of the memory allocation state when selecting the allocable memory location to allocate memory to the current tensor for each allocable memory location corresponding to the current tensor.

[0140] The memory allocation submodule 403 is used to allocate memory to the current tensor based on the allocable memory location corresponding to the memory allocation state with the highest state value score, and update the current memory allocation state of the currently available memory.

[0141] In some embodiments, the location determination submodule 401 is used to: determine all free memory locations in the currently available memory under the current memory allocation state; remove free memory locations from all free memory locations that do not match the tensor type of the current tensor; and determine the allocatable memory location corresponding to the current tensor from the remaining free memory locations based on the tensor size, memory allocation time, memory release time of the current tensor, and the tensor size, memory allocation time, and memory release time of other tensors in the context information.

[0142] In some embodiments, the allocation module further includes a checking module (not shown in the figure) and a memory release module (not shown in the figure). The checking module is used to check whether the tensor of the allocated memory location meets the memory release condition under the current memory allocation state. The release module is used to release the memory location corresponding to the tensor to the currently available memory when the tensor of the allocated memory location meets the memory release condition. The memory release condition includes that the memory release time of the tensor of the allocated memory location is equal to the memory allocation time of the current tensor.

[0143] In some embodiments, the value determination submodule 402 is configured to: when the current iteration is the first iteration, determine the state value score of the memory allocation state when allocating memory to the current tensor at the selected allocable memory location based on the initialized state value score; when the current iteration is not the first iteration, determine the state value score of the memory allocation state when allocating memory to the current tensor at the selected allocable memory location based on the memory allocation information corresponding to each tensor recorded in the historical iterations; wherein, the memory allocation information corresponding to the tensor includes the memory allocation state corresponding to the tensor and the corresponding state value score.

[0144] Figure 10 This is a schematic diagram of another memory allocation device provided in an embodiment of the present disclosure, as shown below. Figure 10 As shown, in some embodiments, the allocation module further includes a value update submodule 404, which is used to: obtain the state penalty value corresponding to each allocated tensor, wherein each allocated tensor includes the current tensor and tensors that were allocated memory locations before the current tensor was allocated; and update the state value score of the memory allocation state corresponding to each allocated tensor according to the state penalty value corresponding to each allocated tensor.

[0145] In some embodiments, such as Figure 10As shown, the allocation module also includes a recording submodule 405, which is used to record the memory allocation information corresponding to the current tensor, wherein the memory allocation information includes the current memory allocation status and the corresponding status value score.

[0146] Figure 11 This is a schematic diagram of another memory allocation device provided in an embodiment of the present disclosure, as shown below. Figure 11 As shown, in some embodiments, the memory allocation device 400 further includes an initialization module 406. The initialization module 406 is used to initialize the memory allocation state of the available memory of the current computational core in the event that memory allocation for the current tensor fails in the current iteration. The allocation module is used to perform the next iteration based on pre-recorded memory allocation information corresponding to each tensor when memory allocation for the current tensor fails in the current iteration. The memory allocation information corresponding to the tensor includes the memory allocation state of the tensor and the corresponding state value score.

[0147] In this embodiment of the disclosure, the memory allocation device 400 can be integrated into a compiler for compiling neural network computation graphs, or into a many-core chip, or into the computing core of a many-core chip, or into an external server or other electronic device. This embodiment of the disclosure does not impose any special limitations on this.

[0148] For a detailed description of each functional module of the memory allocation device 400, please refer to the relevant description in the memory allocation method above, which will not be repeated here.

[0149] Figure 12 This is a block diagram of an electronic device provided in an embodiment of the present disclosure, with reference to... Figure 12 This disclosure provides an electronic device 500, which includes: at least one processor 501; and a memory 502 communicatively connected to the at least one processor 501; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 to enable the at least one processor 501 to execute the memory allocation method described in any of the above embodiments.

[0150] Furthermore, this disclosure also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the memory allocation method of any of the above embodiments.

[0151] This disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the memory allocation method described in any of the above embodiments.

[0152] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0153] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A memory allocation method based on reinforcement learning, applied to a many-core chip, wherein the many-core chip includes multiple computing cores, each computing core being configured with independent memory, the method comprising: In the current iteration, memory allocation operations are performed sequentially on each operator corresponding to the current computational core. Each operator includes at least one tensor, and the tensor type of the tensor is the same as the operator type of the operator it belongs to. The memory allocation operation includes: Given the current memory allocation state of available memory, the allocatable memory location corresponding to the current tensor is determined based on the tensor attribute information and the context information of the current tensor. The current memory allocation state represents the memory allocation status of each tensor allocated in the current iteration. The context information includes the tensor attribute information of other tensors whose memory allocation time is after the current tensor. The tensor attribute information includes tensor type, tensor size, tensor memory allocation time, and memory release time. Based on the state value score of the current tensor at each allocatable memory location, allocate memory for the current tensor and update the current memory allocation state of the currently available memory. In the event that memory allocation for the current tensor fails, the method further includes: Obtain the state penalty value corresponding to each allocated tensor; the allocated tensors include the current tensor and the tensors that were allocated memory locations before the current tensor was allocated. Update the state value score of the memory allocation state corresponding to each allocated tensor based on the state penalty value of each allocated tensor. The step of updating the state value score of the memory allocation state corresponding to each allocated tensor based on the state penalty value corresponding to each allocated tensor includes: Based on the state penalty value corresponding to each allocated tensor, the state value score of the memory allocation state corresponding to each allocated tensor is updated using a predetermined value function. The predetermined value function includes: state(n)_value1 = state(n)_value0 - Sn; Where state(n)_value0 represents the current state value score of the memory allocation state corresponding to the nth tensor, state(n)_value1 represents the updated state value score of the memory allocation state corresponding to the nth tensor, and Sn represents the state penalty value corresponding to the nth tensor.

2. The memory allocation method according to claim 1, wherein allocating memory to the current tensor based on the state value score of the memory allocation state corresponding to each allocatable memory location of the current tensor, and updating the current memory allocation state of the currently available memory, includes: For each allocatable memory location corresponding to the current tensor, determine the state value score of the memory allocation state when selecting that allocatable memory location to allocate memory to the current tensor. Based on the allocable memory location corresponding to the memory allocation state with the highest state value score, allocate memory for the current tensor and update the current memory allocation state of the currently available memory.

3. The memory allocation method according to claim 1, wherein determining the allocatable memory location corresponding to the current tensor based on the tensor attribute information and the context information of the current tensor includes: Determine all free memory locations in the currently available memory under the current memory allocation state; Remove all free memory locations that do not match the tensor type of the current tensor. Based on the current tensor's size, memory allocation time, and memory release time, as well as the tensor sizes, memory allocation times, and memory release times of other tensors in the context information, the allocatable memory location corresponding to the current tensor is determined from the remaining free memory locations.

4. The memory allocation method according to claim 3, wherein determining the allocatable memory location corresponding to the current tensor from the remaining free memory locations based on the current tensor's tensor size, memory allocation time, memory release time, and the tensor sizes, memory allocation times, and memory release times of other tensors in the context information includes: Based on the current tensor's size, memory allocation time, and memory release time, as well as the tensor's size, memory allocation time, and memory release time of other tensors in the context information, memory allocation combinations are performed on the current tensor and other tensors. Each memory allocation combination corresponds to an allocation situation of the current tensor and other tensors in the remaining free memory locations. In memory allocation combination, when any two tensors meet the memory reuse condition, the two tensors are set to be adjacent, and the memory location of the tensor that releases memory first is set to be allocated to the other tensor after the memory is released. For each memory allocation combination, when the memory size required by the memory allocation combination is less than or equal to the total memory size corresponding to the remaining free memory locations, the free memory location corresponding to the current tensor in the memory allocation combination is determined as the allocatable memory location corresponding to the current tensor.

5. The memory allocation method according to claim 1, wherein before determining the allocatable memory location corresponding to the current tensor based on the tensor attribute information and the context information of the current tensor, the method further includes: In the current memory allocation state, check whether the tensor of the allocated memory location meets the memory release condition; When a tensor with an allocated memory location meets the memory release condition, the memory location corresponding to that tensor is released to the currently available memory. The memory release condition includes the fact that the memory release time of a tensor with allocated memory locations is equal to the memory allocation time of the current tensor.

6. The memory allocation method according to claim 1, wherein obtaining the state penalty value corresponding to each allocated tensor includes: Get the distance parameters between each of the allocated tensors and the current tensor; The distance parameter represents the distance between the tensors already allocated in the memory allocation order and the current tensor; For each allocated tensor, the state penalty value corresponding to the allocated tensor is determined based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality. The larger the distance parameter, the smaller the corresponding state penalty value.

7. The memory allocation method according to claim 6, wherein determining the state penalty value corresponding to the allocated tensor based on the distance parameter between the allocated tensor and the current tensor and the penalty cardinality includes: Based on the distance parameter between the allocated tensor and the current tensor and the penalty base, the state penalty value corresponding to the allocated tensor is obtained using the state penalty value function. The state penalty value function includes: Sn = func(Dn, C); Where Sn represents the state penalty value corresponding to the nth allocated tensor, Dn is the distance parameter between the nth allocated tensor and the current tensor, C is the penalty cardinality, func() represents the state penalty value function related to Dn and C, and func() is C raised to the power of Dn.

8. The memory allocation method according to claim 2, wherein after updating the current memory allocation state, the method further includes: Record the memory allocation information corresponding to the current tensor, including the current memory allocation status and the corresponding status value score.

9. The memory allocation method according to claim 2, wherein if memory allocation for the current tensor fails, the method further includes: Initialize the memory allocation state of the available memory for the current computing core; Based on the pre-recorded memory allocation information corresponding to each tensor, the next iteration is performed. The memory allocation information corresponding to the tensor includes the memory allocation state of the tensor and the corresponding state value score.

10. The memory allocation method according to claim 2, wherein determining the state value score of the memory allocation state when selecting the allocatable memory location to allocate memory to the current tensor includes: When the current iteration is the first iteration, the state value score of the memory allocation state is determined based on the initialized state value score when allocating memory to the current tensor at the selected allocable memory location; When the current iteration is not the first iteration, the state value score of the memory allocation state is determined based on the memory allocation information of each tensor recorded in the historical iterations, when the memory is allocated to the current tensor at the selected allocable memory location. The memory allocation information corresponding to the tensor includes the memory allocation status of the tensor and the corresponding status value score.

11. The memory allocation method according to claim 10, wherein determining the state value score of the memory allocation state when allocating memory to the current tensor at the selected allocable memory location, based on the memory allocation information corresponding to each tensor recorded in historical iterations, includes: When there is a memory allocation state in the recorded memory allocation information that is the same as the memory allocation state corresponding to the current tensor at the allocable memory location, determine the state value score corresponding to the memory allocation state in the memory allocation information. If there is no memory allocation state in the recorded memory allocation information that is the same as the memory allocation state corresponding to the current tensor at the allocable memory location, the state value score of the memory allocation state corresponding to the current tensor at the allocable memory location is determined according to the initialized state value score.

12. A memory allocation device based on reinforcement learning, applied to a many-core chip, the many-core chip comprising multiple computing cores, each computing core being configured with independent memory, the device comprising: The allocation module is configured to perform memory allocation operations on each operator corresponding to the current computing core in the current iteration. The operator includes at least one tensor, and the tensor type of the tensor is the same as the operator type of the operator. The allocation module includes: The location determination submodule is used to determine the allocatable memory location corresponding to the current tensor in the current available memory based on the tensor attribute information and the context information of the current tensor, given the current memory allocation state of the currently available memory. The current memory allocation state represents the memory allocation status of each tensor that has been allocated in the current iteration. The context information includes the tensor attribute information of other tensors whose memory allocation time is after that of the current tensor. The tensor attribute information includes tensor type, tensor size, tensor memory allocation time, and memory release time. The memory allocation submodule is used to allocate memory to the current tensor based on the state value score of the memory allocation state corresponding to each allocable memory location of the current tensor, and update the current memory allocation state of the currently available memory. A value update submodule, in the event of a failure to allocate memory to the current tensor, is used to obtain the state penalty value corresponding to each allocated tensor; the allocated tensors include the current tensor and the tensors whose memory locations were allocated before the current tensor was allocated; based on the state penalty value corresponding to each allocated tensor, the state value score of the memory allocation state corresponding to each allocated tensor is updated; based on the state penalty value corresponding to each allocated tensor, the state value score of the memory allocation state corresponding to each allocated tensor is updated using a predetermined value function; wherein the predetermined value function includes: state(n)_value1 = state(n)_value0 - Sn; where state(n)_value0 represents the current state value score of the memory allocation state corresponding to the nth tensor, state(n)_value1 represents the updated state value score of the memory allocation state corresponding to the nth tensor, and Sn represents the state penalty value corresponding to the nth tensor.

13. An electronic device, comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the memory allocation method according to any one of claims 1-11.

14. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by the processor, it implements the memory allocation method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Storage management method and device for computation module

    CN105138289A

  • Memory allocation method and device for neural network

    CN112084038A