Memory Management Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium

Tensor memory allocation information obtained by training solves the problem of low efficiency of large-block memory management in deep learning tasks, and realizes system performance improvement and resource optimization.

CN112559165BActive Publication Date: 2025-05-30ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910913372.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-25
Publication Date
2025-05-30
Estimated Expiration
2039-09-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage large chunks of memory in deep learning tasks, resulting in system performance degradation and waste of resources.

Method used

By obtaining statistical information of tensor application and release requests in the training data, memory allocation information is obtained based on this information, and tensor memory allocation is performed based on this information.

Benefits of technology

It realizes memory allocation with lower system overhead in deep learning tasks, improves memory utilization and reduces resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112559165B_ABST
    Figure CN112559165B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a memory management method, apparatus, electronic device, and computer-readable storage medium. The method includes: obtaining tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time; training memory allocation information according to the tensor memory statistical information; and performing tensor memory allocation according to the memory allocation information. This technical solution is applicable not only to the allocation and management of small chunks of memory, but also to the allocation and management of large chunks of memory. It can not only perform memory allocation with relatively low system overhead, but also effectively improve the utilization rate of memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of memory management, and particularly to a memory management method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] With the development of data technology, deep learning has been widely applied in various fields, such as images, speech, advertising, search, recommendation, natural language processing, and so on. Researchers have found that memory management in deep learning frameworks is a key factor restricting the performance of deep learning. Currently, the optimization research on memory management mainly focuses on GPU video memory, but in fact, the optimization of CPU cluster memory management is also crucial. This is because model training in CPU clusters has the characteristics of large scale and long training time, and the training process requires a large amount of CPU and memory resources. Researchers found that a large amount of CPU resources were wasted on page faults during performance evaluation, which in turn led to memory waste.

[0003] In the prior art, memory management solutions for CPUs include general Malloc library memory management solutions, combined memory management solutions of BFCAllocator and Malloc library, and so on. Although these memory management solutions can improve the performance of network applications to a certain extent, they are not suitable for deep learning tasks. This is because deep learning tasks usually require frequent memory allocation and release. Especially for memory blocks of MB level, the total amount of operations may reach dozens of GB or even hundreds of GB. General memory management solutions generally only manage relatively small memory blocks, such as memory less than 32KB. For large memory blocks, no caching design is made, and they are only allocated in the operating system on demand. In this way, the page fault overhead caused by allocation in a frequently allocated scenario will seriously affect the system performance. And if large memory blocks are cached, it may cause them to be unable to be reused in a relatively long period of time, which in turn will cause relatively large resource waste. Especially for cloud environments, the occupation of a large amount of memory means that the resources available to other applications running on the same machine will be significantly reduced, thus bringing mutual influence between applications. Therefore, there is an urgent need for a memory management method suitable for large memory blocks, which can not only allocate memory with low system overhead, but also effectively improve the utilization rate of memory. Summary of the Invention

[0004] The embodiments of the present invention provide a memory management method, apparatus, electronic device, and computer-readable storage medium.

[0005] In a first aspect, a memory management method is provided in the embodiments of the present invention.

[0006] Specifically, the memory management method includes:

[0007] Obtain tensor application and release requests in the training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0008] Train memory allocation information based on the tensor memory statistical information;

[0009] Perform tensor memory allocation according to the memory allocation information.

[0010] Combined with the first aspect, in the first implementation manner of the first aspect of the embodiments of the present invention, the training of the memory allocation information based on the tensor memory statistical information includes:

[0011] Determine a preset traversal order;

[0012] Traverse the tensor memory statistical information according to the preset traversal order, and train memory allocation information based on the tensor memory statistical information.

[0013] Combined with the first aspect and the first implementation manner of the first aspect, in the second implementation manner of the first aspect of the embodiments of the present invention, the traversing the tensor memory statistical information according to the preset traversal order and training memory allocation information based on the tensor memory statistical information is implemented as:

[0014] Determine target tensor information according to the preset traversal order, where the target tensor information includes target tensor size, target tensor application time, and target tensor release time;

[0015] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks;

[0016] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0017] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0018] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block in ascending order. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order;

[0019] In response to the end of traversing the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0020] Combining the first aspect, the first implementation manner of the first aspect, and the second implementation manner of the first aspect, in the third implementation manner of the first aspect of the present disclosure, the determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is implemented as:

[0021] Determine the memory address interval of the memory address linked list block;

[0022] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0023] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0024] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0025] Combining the first implementation manner, the second implementation manner, the third implementation manner, and the fourth implementation manner of the first aspect, in the fourth implementation manner of the first aspect of the present disclosure, the performing tensor memory allocation according to the memory allocation information includes:

[0026] Initialize the memory pool allocator according to the memory allocation information;

[0027] Perform tensor memory allocation based on the initialized memory pool allocator.

[0028] Combining the first aspect, the first implementation manner, the second implementation manner, the third implementation manner, and the fourth implementation manner of the first aspect, in the fifth implementation manner of the first aspect of the present disclosure, it further includes:

[0029] In response to detecting a tensor memory allocation failure event, reallocate the memory for the memory reallocation failure tensor according to a preset memory allocation rule.

[0030] Combined with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, and the fifth implementation manner of the first aspect, in the sixth implementation manner of the first aspect of the present disclosure, the reallocating the memory for the memory allocation failure tensor according to a preset memory allocation rule in response to detecting a tensor memory allocation failure event includes:

[0031] In response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failure tensor according to the size of the memory allocation failure tensor;

[0032] Determine whether there is a free cache block in the candidate memory address linked list block;

[0033] When there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill the first header information for the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0034] When there is no free cache block in the candidate memory address linked list block, apply for memory allocation from the operating system based on the size of the memory allocation failure tensor, and fill the second header information for the allocated memory, where the second header information stores the operating system memory allocation information.

[0035] Combined with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, and the seventh implementation manner of the first aspect, in the eighth implementation manner of the first aspect of the present disclosure, it further includes:

[0036] In response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

[0037] Combined with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, and the seventh implementation manner of the first aspect, in the eighth implementation manner of the first aspect of the present disclosure, the releasing the memory according to a preset memory release rule in response to detecting a tensor memory release command includes:

[0038] In response to detecting a tensor memory release command, obtain header information corresponding to the tensor memory release command;

[0039] Release the memory according to the header information.

[0040] In a second aspect, an embodiment of the present invention provides a memory management method.

[0041] Specifically, the memory management method includes:

[0042] Obtain tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0043] Traverse the tensor memory statistical information, and insert the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size to obtain memory allocation information;

[0044] Initialize a memory pool allocator according to the memory allocation information, and perform tensor memory allocation based on the initialized memory pool allocator.

[0045] Combined with the second aspect, in a first implementation manner of the second aspect of the embodiment of the present invention, the traversing the tensor memory statistical information, inserting the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size to obtain memory allocation information is implemented as:

[0046] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0047] Determine a target memory address linked list block corresponding to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache block;

[0048] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0049] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0050] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block in ascending order. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order;

[0051] In response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0052] Combining the second aspect and the first implementation manner of the second aspect, in the second implementation manner of the second aspect of the embodiments of the present invention, the determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is implemented as:

[0053] Determine the memory address interval of the memory address linked list block;

[0054] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0055] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0056] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0057] Combining the second aspect, the first implementation manner of the second aspect and the second implementation manner of the second aspect, in the third implementation manner of the second aspect of the present disclosure, it further includes:

[0058] In response to detecting a tensor memory allocation failure event, reallocate the memory reallocation failure tensor according to a preset memory allocation rule.

[0059] Combining the first implementation manner, the second implementation manner and the third implementation manner of the second aspect, in the fourth implementation manner of the second aspect of the present disclosure, the reallocating the memory allocation failure tensor according to a preset memory allocation rule in response to detecting a tensor memory allocation failure event includes:

[0060] In response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the memory allocation failure tensor size according to the memory allocation failure tensor size;

[0061] Determine whether there is a free cache block in the candidate memory address linked list block;

[0062] When there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill the first header information into the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0063] When there is no free cache block in the candidate memory address linked list block, apply for memory allocation from the operating system based on the memory allocation failure tensor size, and fill the second header information into the allocated memory, where the second header information stores the operating system memory allocation information.

[0064] Combined with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, and the fourth implementation manner of the second aspect, in the fifth implementation manner of the second aspect of the present disclosure, it further includes:

[0065] In response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

[0066] Combined with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, the fourth implementation manner of the second aspect, and the fifth implementation manner of the second aspect, in the sixth implementation manner of the second aspect of the present disclosure, the step of releasing the memory according to a preset memory release rule in response to detecting a tensor memory release command includes:

[0067] In response to detecting a tensor memory release command, obtain the header information corresponding to the tensor memory release command;

[0068] Release the memory according to the header information.

[0069] In a third aspect, an embodiment of the present invention provides a memory management method.

[0070] Specifically, the memory management method includes:

[0071] Obtain tensor memory statistical information, where the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0072] Traverse the tensor memory statistics and insert the tensor memory statistics into the cache block of the target memory address linked list block corresponding to the tensor size;

[0073] In response to the end of traversing the tensor memory statistics, determine the current memory address linked list block information as the memory allocation information.

[0074] Combined with the third aspect, in the first implementation manner of the third aspect of the embodiments of the present invention, the traversing the tensor memory statistics and inserting the tensor memory statistics into the cache block of the target memory address linked list block corresponding to the tensor size is implemented as:

[0075] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0076] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistics are stored in the cache block;

[0077] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0078] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0079] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order.

[0080] Combined with the third aspect and the first implementation manner of the third aspect, in the second implementation manner of the third aspect of the embodiments of the present invention, the determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is implemented as:

[0081] Determine the memory address interval of the memory address linked list block;

[0082] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0083] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0084] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0085] Combined with the third aspect, the first implementation manner of the third aspect and the second implementation manner of the third aspect, in the third implementation manner of the third aspect of the present disclosure, it further includes:

[0086] In response to detecting a tensor memory allocation failure event, reallocate the memory for the memory reallocation failure tensor according to a preset memory allocation rule.

[0087] Combined with the first implementation manner, the second implementation manner and the third implementation manner of the third aspect, in the fourth implementation manner of the third aspect of the present disclosure, the reallocating the memory for the memory allocation failure tensor according to a preset memory allocation rule in response to detecting a tensor memory allocation failure event includes:

[0088] In response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the memory allocation failure tensor size according to the memory allocation failure tensor size;

[0089] Determine whether there is a free cache block in the candidate memory address linked list block;

[0090] When there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill the first header information for the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0091] When there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the memory allocation failure tensor size, and fill the second header information for the allocated memory, where the second header information stores the operating system memory allocation information.

[0092] Fourthly, an embodiment of the present invention provides a memory management device.

[0093] Specifically, the memory management device includes:

[0094] A first acquisition module, configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0095] A training module, configured to train memory allocation information according to the tensor memory statistical information;

[0096] A first allocation module, configured to allocate tensor memory according to the memory allocation information.

[0097] Combined with the fourth aspect, in the first implementation manner of the fourth aspect of the embodiment of the present invention, the training module includes:

[0098] A first determination sub-module, configured to determine a preset traversal order;

[0099] A training sub-module, configured to traverse the tensor memory statistical information according to the preset traversal order, and train memory allocation information according to the tensor memory statistical information.

[0100] Combined with the fourth aspect and the first implementation manner of the fourth aspect, in the second implementation manner of the fourth aspect of the embodiment of the present invention, the training sub-module is configured to:

[0101] Determine target tensor information according to the preset traversal order, where the target tensor information includes target tensor size, target tensor application time, and target tensor release time;

[0102] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks;

[0103] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0104] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0105] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block in ascending order. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order;

[0106] In response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0107] Combined with the fourth aspect, the first implementation manner of the fourth aspect, and the second implementation manner of the fourth aspect, in the third implementation manner of the fourth aspect of the present disclosure, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is configured to:

[0108] Determine the memory address interval of the memory address linked list block;

[0109] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0110] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0111] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0112] Combined with the fourth aspect, the first implementation manner of the fourth aspect, the second implementation manner of the fourth aspect, and the third implementation manner of the fourth aspect, in the fourth implementation manner of the fourth aspect of the present disclosure, the first allocation module includes:

[0113] An initialization sub-module, configured to initialize the memory pool allocator according to the memory allocation information;

[0114] A first allocation sub-module, configured to perform tensor memory allocation based on the initialized memory pool allocator.

[0115] Combined with the fourth aspect, the first implementation manner of the fourth aspect, the second implementation manner of the fourth aspect, the third implementation manner of the fourth aspect, and the fourth implementation manner of the fourth aspect, in the fifth implementation manner of the fourth aspect of the present disclosure, it further includes:

[0116] A first redistribution module, configured to, in response to detecting a tensor memory allocation failure event, redistribute the memory allocation failure tensor according to a preset memory allocation rule.

[0117] Combined with the fourth aspect, the first implementation manner of the fourth aspect, the second implementation manner of the fourth aspect, the third implementation manner of the fourth aspect, the fourth implementation manner of the fourth aspect, and the fifth implementation manner of the fourth aspect, in the sixth implementation manner of the fourth aspect of the present disclosure, the first redistribution module includes:

[0118] A second determination sub-module, configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failure tensor according to the size of the memory allocation failure tensor;

[0119] A third determination sub-module, configured to determine whether there is a free cache block in the candidate memory address linked list block;

[0120] A second allocation sub-module, configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill the first header information in the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0121] A third allocation sub-module, configured to, when there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the memory allocation failure tensor, and fill the second header information in the applied memory, where the second header information stores the operating system memory allocation information.

[0122] Combined with the fourth aspect, the first implementation manner of the fourth aspect, the second implementation manner of the fourth aspect, the third implementation manner of the fourth aspect, the fourth implementation manner of the fourth aspect, the fifth implementation manner of the fourth aspect, and the sixth implementation manner of the fourth aspect, in the seventh implementation manner of the fourth aspect of the present disclosure, it further includes:

[0123] A first release module, configured to, in response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

[0124] Combined with the fourth aspect, the first implementation manner of the fourth aspect, the second implementation manner of the fourth aspect, the third implementation manner of the fourth aspect, the fourth implementation manner of the fourth aspect, the fifth implementation manner of the fourth aspect, the sixth implementation manner of the fourth aspect, and the seventh implementation manner of the fourth aspect, in the eighth implementation manner of the fourth aspect of the present disclosure, the first release module includes:

[0125] A first acquisition sub-module, configured to acquire header information corresponding to the tensor memory release command in response to detecting the tensor memory release command;

[0126] A first release sub-module, configured to release the memory according to the header information.

[0127] Fifth aspect, an embodiment of the present invention provides a memory management device.

[0128] Specifically, the memory management device includes:

[0129] A second acquisition module, configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0130] A first insertion module, configured to traverse the tensor memory statistical information and insert the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size to obtain memory allocation information;

[0131] A second allocation module, configured to initialize a memory pool allocator according to the memory allocation information and perform tensor memory allocation based on the initialized memory pool allocator.

[0132] Combined with the fifth aspect, in the first implementation manner of the fifth aspect of the embodiment of the present invention, the first insertion module is configured to:

[0133] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0134] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache block;

[0135] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0136] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order;

[0137] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order;

[0138] In response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0139] Combining the fifth aspect and the first implementation manner of the fifth aspect, in the second implementation manner of the fifth aspect of the embodiments of the present invention, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is configured to:

[0140] Determine the memory address interval of the memory address linked list block;

[0141] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0142] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0143] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0144] Combining the fifth aspect, the first implementation manner of the fifth aspect and the second implementation manner of the fifth aspect, in the third implementation manner of the fifth aspect of the present disclosure, it further includes:

[0145] A second reallocation module, configured to reallocate the memory reallocation failed tensors according to a preset memory allocation rule in response to detecting a tensor memory allocation failure event.

[0146] In combination with the fifth aspect, the first implementation manner of the fifth aspect, the second implementation manner of the fifth aspect, and the third implementation manner of the fifth aspect, in the fourth implementation manner of the fifth aspect of the present disclosure, the second redistribution module includes:

[0147] A fourth determination sub-module, configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the tensor size of the memory allocation failure according to the tensor size of the memory allocation failure;

[0148] A fifth determination sub-module, configured to determine whether there is a free cache block in the candidate memory address linked list block;

[0149] A fourth allocation sub-module, configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the tensor of the memory allocation failure, and fill a first header information for the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0150] A fifth allocation sub-module, configured to, when there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the tensor size of the memory allocation failure, and fill a second header information for the obtained memory, where the second header information stores the operating system memory allocation information.

[0151] In combination with the fifth aspect, the first implementation manner of the fifth aspect, the second implementation manner of the fifth aspect, the third implementation manner of the fifth aspect, and the fourth implementation manner of the fifth aspect, in the fifth implementation manner of the fifth aspect of the present disclosure, it further includes:

[0152] A second release module, configured to, in response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

[0153] In combination with the fifth aspect, the first implementation manner of the fifth aspect, the second implementation manner of the fifth aspect, the third implementation manner of the fifth aspect, the fourth implementation manner of the fifth aspect, and the fifth implementation manner of the fifth aspect, in the sixth implementation manner of the fifth aspect of the present disclosure, the second release module includes:

[0154] A second acquisition sub-module, configured to, in response to detecting a tensor memory release command, acquire header information corresponding to the tensor memory release command;

[0155] A second release sub-module, configured to release the memory according to the header information.

[0156] Sixth aspect, an embodiment of the present invention provides a memory management device.

[0157] Specifically, the memory management device includes:

[0158] A third acquisition module, configured to acquire tensor memory statistical information, where the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0159] A second insertion module, configured to traverse the tensor memory statistical information and insert the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size;

[0160] A determination module, configured to, in response to the end of traversing the tensor memory statistical information, determine the current memory address linked list block information as the memory allocation information.

[0161] Combined with the sixth aspect, in the first implementation manner of the sixth aspect of the embodiments of the present invention, the second insertion module is configured to:

[0162] Determine a preset traversal order and determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0163] Determine a target memory address linked list block corresponding to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache block;

[0164] Determine whether there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information;

[0165] When there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not generate a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0166] When there is no cache block in the target memory address linked list block that does not have a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block from smallest to largest. When there is a cache block in the memory address linked list block that does not have a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not have a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not have a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order.

[0167] Combined with the sixth aspect and the first implementation manner of the sixth aspect, in the second implementation manner of the sixth aspect of the embodiments of the present invention, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is configured as:

[0168] Determine the memory address interval of the memory address linked list block;

[0169] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0170] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0171] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0172] Combined with the sixth aspect, the first implementation manner of the sixth aspect, the second implementation manner of the sixth aspect, in the third implementation manner of the sixth aspect of the present disclosure, it further includes:

[0173] A third reallocation module, configured to reallocate the memory reallocation failed tensor according to a preset memory allocation rule in response to detecting a tensor memory allocation failure event.

[0174] Combined with the sixth aspect, the first implementation manner of the sixth aspect, the second implementation manner of the sixth aspect, and the third implementation manner of the sixth aspect, in the fourth implementation manner of the sixth aspect of the present disclosure, the third reallocation module includes:

[0175] A sixth determination sub-module, configured to determine a candidate memory address linked list block corresponding to the memory allocation failed tensor size according to the memory allocation failed tensor size in response to detecting a tensor memory allocation failure event;

[0176] The seventh determination sub-module is configured to determine whether there is a free cache block in the candidate memory address linked list block;

[0177] The sixth allocation sub-module is configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failed tensor, and fill the first header information into the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0178] The seventh allocation sub-module is configured to, when there is no free cache block in the candidate memory address linked list block, apply for memory allocation from the operating system based on the size of the memory allocation failed tensor, and fill the second header information into the allocated memory, where the second header information stores the operating system memory allocation information.

[0179] In a seventh aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. The memory is used to store one or more computer instructions for supporting the memory management device to execute the above-mentioned memory management method, and the processor is configured to execute the computer instructions stored in the memory. The memory management device may further include a communication interface for communicating between the memory management device and other devices or communication networks.

[0180] In an eighth aspect, an embodiment of the present invention provides a computer-readable storage medium for storing computer instructions used by the memory management device, which includes computer instructions for executing the above-mentioned memory management method related to the memory management device.

[0181] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0182] The above technical solution trains memory allocation information based on representative training data, and then performs subsequent tensor memory allocation according to the memory allocation information. This technical solution is applicable not only to the allocation and management of small memory blocks, but also to the allocation and management of large memory blocks. It can not only perform memory allocation with low system overhead, but also effectively improve the utilization rate of memory.

[0183] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0184] Combined with the drawings, through the following detailed description of non-limiting embodiments, other features, objects, and advantages of the embodiments of the present invention will become more obvious. In the drawings:

[0185] Figure 1 A flowchart showing a memory management method according to an embodiment of the present invention is shown;

[0186] Figure 2 Show the steps S102 of the memory management method according to Figure 1 The flowchart shown in the illustrated embodiment;

[0187] Figure 3 Show the steps according to Figure 1 The flowchart of the training memory allocation information shown in the illustrated embodiment;

[0188] Figure 4 Show the steps according to Figure 1 The flowchart of step S103 of the memory management method shown in the illustrated embodiment;

[0189] Figure 5 Show the flowchart of the memory management method according to another embodiment of the present invention;

[0190] Figure 6 Show the steps according to Figure 5 The flowchart of step S504 of the memory management method shown in the illustrated embodiment;

[0191] Figure 7 Show the steps according to Figure 1 The flowchart of the tensor reallocation shown in the illustrated embodiment;

[0192] Figure 8 Show the flowchart of the memory management method according to another embodiment of the present invention;

[0193] Figure 9 Show the steps according to Figure 8 The flowchart of step S805 of the memory management method shown in the illustrated embodiment;

[0194] Figure 10 Show the steps according to Figure 1 The flowchart of the tensor release shown in the illustrated embodiment;

[0195] Figure 11 Show the schematic diagram of the memory management application scenario according to an embodiment of the present invention;

[0196] Figure 12 Show the flowchart of the memory management method according to another embodiment of the present invention;

[0197] Figure 13 Show the flowchart of the memory management method according to still another embodiment of the present invention;

[0198] Figure 14 Show the structural block diagram of the memory management device according to an embodiment of the present invention;

[0199] Figure 15 Show the steps according to Figure 14Block diagram of the training module 1402 of the memory management device according to the illustrated embodiment;

[0200] Figure 16 Illustrating according to Figure 14 Block diagram of the first allocation module 1403 of the memory management device according to the illustrated embodiment;

[0201] Figure 17 Block diagram of the memory management device according to another embodiment of the present invention;

[0202] Figure 18 Illustrating according to Figure 17 Block diagram of the first reallocation module 1704 of the memory management device according to the illustrated embodiment;

[0203] Figure 19 Block diagram of the memory management device according to another embodiment of the present invention;

[0204] Figure 20 Illustrating according to Figure 19 Block diagram of the first release module 1905 of the memory management device according to the illustrated embodiment;

[0205] Figure 21 Block diagram of the memory management device according to another embodiment of the present invention;

[0206] Figure 22 Block diagram of the memory management device according to yet another embodiment of the present invention;

[0207] Figure 23 Block diagram of the electronic device according to an embodiment of the present invention;

[0208] Figure 24 It is a schematic structural diagram of a computer system suitable for implementing the memory management method according to an embodiment of the present invention. Detailed implementation manners

[0209] Hereinafter, exemplary embodiments of the embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for clarity, parts irrelevant to the description of the exemplary embodiments are omitted in the drawings.

[0210] In the embodiments of the present invention, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0211] In addition, it should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The embodiments of the present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0212] The technical solution provided by the embodiment of the present invention trains memory allocation information based on representative training data, and then performs subsequent tensor memory allocation according to the memory allocation information. This technical solution is applicable not only to the allocation and management of small pieces of memory, but also to the allocation and management of large pieces of memory. It can not only perform memory allocation with low system overhead, but also effectively improve the utilization rate of memory.

[0213] Figure 1 The flowchart showing a memory management method according to an embodiment of the present invention is as Figure 1 shown, and the memory management method includes the following steps S101-S103:

[0214] In step S101, tensor application and release requests in the training data are obtained, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0215] In step S102, memory allocation information is trained according to the tensor memory statistical information;

[0216] In step S103, tensor memory allocation is performed according to the memory allocation information.

[0217] As mentioned above, with the development of data technology, deep learning has been widely applied in various fields, such as images, speech, advertising, search, recommendation, natural language processing, and so on. Researchers have found that memory management in deep learning frameworks is a key factor restricting the performance of deep learning. At present, the optimization research on memory management mainly focuses on GPU video memory, but in fact, the optimization of CPU cluster memory management is also crucial. This is because model training in CPU clusters has the characteristics of large scale and long training time, and a large amount of CPU and memory resources are required during the training process. Researchers found that a large amount of CPU resources were wasted on page faults during the performance evaluation process, which in turn led to waste of memory.

[0218] In the prior art, the memory management schemes for CPUs include the general Malloc library memory management scheme, the BFCAllocator and Malloc library combined memory management scheme, and so on. Although these memory management schemes can improve the performance of network applications to a certain extent, they are not applicable to deep learning tasks. This is because deep learning tasks usually require frequent memory allocation and deallocation. Especially for memory blocks in the MB level, the total amount of operations may reach dozens of GB or even hundreds of GB. The general memory management scheme generally only manages relatively small memory blocks, such as memory less than 32KB. For large memory blocks, no caching design is made, and they are only allocated in the operating system on demand. In such a frequently allocated scenario, the page fault overhead caused by allocation will seriously affect the system performance. If large memory blocks are cached, it may lead to the situation that they cannot be reused in a relatively long period of time, which will in turn cause relatively large resource waste. Especially in the cloud environment, the occupation of a large amount of memory means that the resources available to other applications running on the same machine will be significantly reduced, thus bringing mutual influence between applications. Therefore, there is an urgent need for a memory management method applicable to large memory blocks, which can not only perform memory allocation with relatively low system overhead, but also effectively improve the utilization rate of memory.

[0219] Considering the above problems, in this embodiment, a memory management method is proposed. This method trains memory allocation information based on representative training data, and then performs subsequent tensor memory allocation according to the memory allocation information. This technical solution is not only applicable to the allocation and management of small memory blocks, but also applicable to the allocation and management of large memory blocks. It can not only perform memory allocation with relatively low system overhead, but also effectively improve the utilization rate of memory.

[0220] Considering that although the cycle sizes and lifecycles of the memory blocks required in a mini-batch of deep learning tasks are different, the mini-batch data has cycle predictability, which makes the cross-mini-batch lifecycle relatively fixed. Among them, the mini-batch data refers to the several small training data sets obtained by dividing the entire training data set when performing the gradient descent algorithm on the training data set in the deep learning scenario. Each time, only one small training data set is trained. Since the gradient direction of the small training data set is not very different from the gradient direction of the entire training data set, the correctness of training can be guaranteed, and at the same time, the problem of huge computational amount caused by all data sets participating in training at one time can be avoided.

[0221] Based on the data characteristics of the above deep learning tasks, in the embodiments of the present invention, memory allocation information is trained based on training data. Since the training data has a certain representativeness and can represent the characteristics of other data, the memory allocation information can be considered to be applicable to other data as well. Subsequently, tensor memory allocation processing can be performed on other data according to the trained memory allocation information.

[0222] That is, in an embodiment of the present invention, the training data refers to the data used to train the memory allocation information that can be used for other data processing. Specifically, the training data can be the first K mini-batch data in a certain deep learning task. Among them, the applications and release requests for tensors generated in the training data actually participate in the training of the memory allocation information. Here, the tensor refers to an array of arbitrary number of dimensions composed of a set of primitive values and is the data unit for storing data in TensorFlow. The tensor application and release requests carry tensor memory statistical information associated with memory allocation and decisive for the memory allocation strategy. The tensor memory statistical information at least includes information such as tensor size, tensor application time, and tensor release time. The tensor size is used to determine the size of the required memory space, and the tensor application time and tensor release time are used to determine the lifetime of the tensor memory statistical information, which can further be used as a basis for judging whether the memory space corresponding to the tensor size is reusable. That is, if there is an overlap in the request lifetimes of a certain memory space, it means that this memory space is not reusable for the overlapping requests. On the contrary, if there is no overlap in the request lifetimes of a certain memory space, it means that this memory space is reusable for the relevant requests.

[0223] In an embodiment of the present invention, the tensor memory statistical information may include one or more tensor memory statistical information. Subsequently, when training the memory allocation information, training will be performed based on all the tensor memory statistical information.

[0224] In an embodiment of the present invention, as Figure 2 shown, the step S102, that is, the step of training the memory allocation information according to the tensor memory statistical information, includes the following steps S201 - S202:

[0225] In step S201, determine a preset traversal order;

[0226] In step S202, traverse the tensor memory statistical information according to the preset traversal order, and train the memory allocation information according to the tensor memory statistical information.

[0227] As mentioned above, when training the memory allocation information according to the tensor memory statistics, all the tensor memory statistics are used for training, that is, all the information in the tensor memory statistics needs to be traversed so that each tensor memory statistic contributes to the generation of the memory allocation information. Specifically, first, a preset traversal order is determined; then, the tensor memory statistics are traversed according to the preset traversal order, and the memory allocation information is trained based on all the tensor memory statistics.

[0228] In an embodiment of the present invention, the preset traversal order can be set according to the needs of actual applications and the data characteristics of the tensor memory statistics, and the present disclosure does not specifically limit it. For example, the preset traversal order can be in the order from largest to smallest tensor size, or of course, in the order from smallest to largest tensor size or in a randomly selected order according to tensor size, etc.

[0229] In an embodiment of the present invention, step S202, that is, the step of traversing the tensor memory statistics according to the preset traversal order and training the memory allocation information according to the tensor memory statistics, can be implemented as:

[0230] Determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0231] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistics are stored in the cache blocks;

[0232] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0233] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0234] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block in ascending order. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order;

[0235] In response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0236] In order to find memory allocation information of appropriate size for the tensor memory statistics information, while maximizing the reuse of memory, improving the utilization rate of memory, and reducing the waste of memory resources, in this embodiment, a traversal and trial method is used to train the memory allocation information based on the tensor memory statistics information.

[0237] Specifically, as Figure 3 shown, the process starts. First, determine the target tensor information that currently requires memory allocation according to the preset traversal order, that is, traverse the tensor information in descending order of tensor size. Among them, the target tensor information at least includes information such as the target tensor size, the target tensor application time, and the target tensor release time.

[0238] Then, determine the target memory address linked list block corresponding to the target tensor size according to the target tensor size. Among them, one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block. Specifically, the size of the cache block should be greater than or equal to the target tensor size to sufficiently meet the memory storage requirements of the target tensor. Among them, the cache block may be an empty cache block or may already store one or more historical tensor memory statistics information. Subsequently, it is determined whether the target tensor information can be inserted into the cache block according to the comparison of the lifetimes between the target tensor information and the historical tensor memory statistics information.

[0239] Then, based on the lifetime comparison mentioned above, it is determined whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information. In an embodiment of the present invention, the preset conflict refers to a lifetime conflict. As mentioned above, if there is an intersection between the lifetime of the target tensor information and the lifetime of the historical tensor memory statistical information already stored in the cache block, it indicates that there is a lifetime conflict between the target tensor information and the cache block, and the memory space stored in this cache block is not reusable for the target tensor information. On the contrary, if there is no intersection between the lifetime of the target tensor information and the lifetime of the historical tensor memory statistical information already stored in the cache block, it indicates that there is no lifetime conflict between the target tensor information and the cache block, and the memory space stored in this cache block is reusable for the target tensor information.

[0240] Therefore, if there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, it means that the memory space stored in this cache block is reusable for the target tensor information. The target tensor information can be inserted into the cache block that does not cause a preset conflict with the target tensor information, and the memory allocation relationship for the current target tensor information is ended. The next target tensor information can be determined according to the preset traversal order to continue the memory allocation.

[0241] If there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, it means that the memory space stored in this cache block is not reusable for the target tensor information and cannot receive the memory allocation request corresponding to the target tensor information. At this time, it is necessary to find a cache block in other memory address linked list blocks that can receive the memory allocation request corresponding to the target tensor information. To improve the search efficiency, in an embodiment of the present invention, within the range of other memory address linked list blocks larger than the target memory address linked list block, traverse the memory address linked list blocks in ascending order of cache block size to find a cache block that can receive the memory allocation request corresponding to the target tensor information. Further, similar to the above description, if there is a cache block in the memory address linked list blocks within the search range that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and create a shared pointer in the target memory address linked list block to point to this cache block, return the traversal process of the tensor information, and determine the next target tensor information for processing according to the preset traversal order. However, if there is no cache block in the memory address linked list blocks that does not cause a preset conflict with the target tensor information, only the target memory address linked list block can be returned, create a new cache block in the target memory address linked list block to receive the memory allocation request corresponding to the target tensor information, insert the target tensor information into the new cache block, and return the traversal process of the tensor information, and determine the next target tensor information for processing according to the preset traversal order.

[0242] Finally, after the traversal of the tensor memory statistics information is completed, the obtained current memory address linked list block information is determined as the memory allocation information, and the memory allocation information can be used to perform memory allocation for other data.

[0243] In this way, not only can the appropriate memory allocation information be found for the tensor memory statistics information, but also the reuse of memory can be maximally realized, the utilization rate of memory can be improved, and the waste of memory resources can be reduced.

[0244] In an embodiment of the present invention, the step of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size may be implemented as:

[0245] Determine the memory address interval of the memory address linked list block;

[0246] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0247] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0248] Determine the memory address linked list block corresponding to the size of the target memory address as the target memory address linked list block.

[0249] In order to make full use of the storage space of the memory address linked list block while meeting the requirements of the target tensor size, in this embodiment, based on the memory address interval of the memory address linked list block and the target tensor size, determine the target memory address linked list block corresponding to the target tensor size. Specifically, first determine the memory address interval of the memory address linked list block, where the memory address interval can be set according to the needs of actual applications, and the present disclosure does not make specific limitations on it; then round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; then multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; finally, determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0250] For example, if the memory address interval is set to 4KB and the target tensor size is 33KB, 33KB / 4KB = 8.25, then the required number of memory address intervals is 9, and the target memory address size is 4KB × 9 = 36KB.

[0251] In an embodiment of the present invention, as Figure 4 shown, step S103, that is, the step of performing tensor memory allocation according to the memory allocation information, includes the following steps S401 - S402:

[0252] In step S401, initialize the memory pool allocator according to the memory allocation information;

[0253] In step S402, perform tensor memory allocation based on the initialized memory pool allocator.

[0254] After obtaining the memory allocation information through training, tensor memory allocation processing can be performed on other data according to the memory allocation information. Specifically, first initialize the memory pool allocator according to the memory allocation information; then perform tensor memory allocation based on the initialized memory pool allocator. For example, if there are two cache blocks in the memory address linked list block with a cache block size of 36KB, then 2 caches with a size of 36KB can be initialized in the cache of the memory pool allocator for memory applications corresponding to all tensors of the size of this memory address linked list block.

[0255] In an embodiment of the present invention, the method further includes the step of re - allocating the tensors with memory allocation failure in accordance with a preset memory allocation rule in response to detecting a tensor memory allocation failure event, that is, asFigure 5 As shown, the memory management method includes the following steps S501 - S504:

[0256] In step S501, obtain the tensor application and release requests in the training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0257] In step S502, train memory allocation information based on the tensor memory statistical information;

[0258] In step S503, perform tensor memory allocation according to the memory allocation information;

[0259] In step S504, in response to detecting a tensor memory allocation failure event, re - allocate the tensor with memory allocation failure according to a preset memory re - allocation rule.

[0260] Considering that some variable sparse tensors may cause the cache blocks determined according to the memory allocation information to be unable to meet their memory requirements, in order to also perform reasonable memory allocation for these special tensors, in this embodiment, if the memory allocation for the above - mentioned preset tensors fails, that is, a tensor memory allocation failure event is detected, then re - allocate the tensor with memory allocation failure according to a preset memory re - allocation rule.

[0261] In an embodiment of the present invention, the preset memory re - allocation rule can be set according to the needs of actual applications and the data characteristics of tensors, and the present disclosure does not make specific limitations on it.

[0262] In an embodiment of the present invention, as Figure 6 shown, step S504, that is, the step of, in response to detecting a tensor memory allocation failure event, re - allocating the tensor with memory allocation failure according to a preset memory re - allocation rule, includes the following steps S601 - S604:

[0263] In step S601, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked - list block corresponding to the size of the tensor with memory allocation failure according to the size of the tensor with memory allocation failure;

[0264] In step S602, determine whether there is an idle cache block in the candidate memory address linked - list block;

[0265] In step S603, when there is an idle cache block in the candidate memory address linked - list block, allocate the idle cache block to the tensor with memory allocation failure, and fill the first header information for the idle cache block, where the first header information stores the information of the memory address linked - list block where the idle cache block is located;

[0266] In step S604, when there is no free cache block in the candidate memory address linked list block, memory allocation is applied to the operating system based on the memory allocation failure tensor size, and the obtained memory is filled with second header information, where the operating system memory allocation information is stored in the second header information.

[0267] In this embodiment, as Figure 7 shown, after detecting a tensor memory allocation failure event, first determine a candidate memory address linked list block corresponding to the memory allocation failure tensor size according to the memory allocation failure tensor size, where the determination method of the candidate memory address linked list block can refer to the determination method of the target memory address linked list block above, and the present disclosure will not elaborate here; then determine whether there is a free cache block in the candidate memory address linked list block; when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill the free cache block with first header information, where the memory address linked list block information where the free cache block is located is stored in the first header information, which is used to identify the storage and memory allocation location of the memory allocation failure tensor and provide a basis for subsequent memory release; when there is no free cache block in the candidate memory address linked list block, directly apply for memory allocation to the operating system based on the memory allocation failure tensor size, and fill the obtained memory with second header information, where the operating system memory allocation information is stored in the second header information, which is used to identify the storage and memory allocation location of the memory allocation failure tensor and provide a basis for subsequent memory release.

[0268] In an embodiment of the present invention, the method further includes a step of releasing memory according to a preset memory release rule in response to detecting a tensor memory release command, that is, as Figure 8 shown, the memory management method includes the following steps S801 - S805:

[0269] In step S801, obtain tensor application and release requests in the training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0270] In step S802, train memory allocation information according to the tensor memory statistical information;

[0271] In step S803, perform tensor memory allocation according to the memory allocation information;

[0272] In step S804, in response to detecting a tensor memory allocation failure event, re - allocate the memory allocation failure tensor according to a preset memory re - allocation rule;

[0273] In step S805, in response to detecting a tensor memory release command, the memory is released according to a preset memory release rule.

[0274] In this embodiment, in order to improve the utilization rate of the memory, after detecting a tensor memory release command, the occupied memory can be released according to a preset memory release rule.

[0275] In an embodiment of the present invention, the preset memory release rule can be set according to the needs of actual applications and the data characteristics of tensors, and the present disclosure does not make specific limitations thereto.

[0276] In an embodiment of the present invention, as Figure 9 shown, step S805, that is, the step of releasing the memory according to a preset memory release rule in response to detecting a tensor memory release command, includes the following steps S901 - S902:

[0277] In step S901, in response to detecting a tensor memory release command, header information corresponding to the tensor memory release command is obtained;

[0278] In step S902, the memory is released according to the header information.

[0279] As mentioned above, the header information is used to identify the storage and memory allocation location of the memory allocation failed tensor. Therefore, as Figure 10 shown, after detecting a tensor memory release command, the header information corresponding to the tensor memory release command can be obtained according to the pointer information of the tensor release, and then the memory can be released according to the header information, so that the memory can return to the idle state in the shortest time to store other information in a timely manner. For example, if the header information indicates that the tensor is stored in a certain cache block of a certain memory address linked list block, the released memory information is returned to the memory address linked list block, otherwise it can be directly returned to the operating system.

[0280] Next, taking an application scenario as an example, the technical solution of the present invention will be further described. In this application scenario, first, training data is determined and the tensor memory statistical information corresponding to the tensor application and release requests in the training data is obtained. Then, all the tensor memory statistical information is traversed and compared with the size and lifetime of the cache blocks in the memory address linked list block until all the tensor memory statistical information is placed in the memory address linked list block. The finally formed memory allocation information is the memory allocation information obtained by training. Subsequently, memory can be allocated for other data based on the memory allocation information. As Figure 11 shown, Figure 11Among them, after traversing all tensor memory statistics and comparing them with the sizes and lifetimes of cache blocks in the memory address linked list block respectively, the obtained memory allocation information is that there are three cache blocks in the memory address linked list block with a cache block size of 36KB, two cache blocks in the memory address linked list block with a cache block size of 40KB, and two cache blocks in the memory address linked list block with a cache block size of 48KB. Then, when initializing the cache of the memory pool allocator subsequently, 3 caches with a size of 36KB can be initialized in the memory pool allocator to correspond to the memory applications of all tensors with the size of this memory address linked list block, 2 caches with a size of 40KB can be initialized to correspond to the memory applications of all tensors with the size of this memory address linked list block, and 2 caches with a size of 48KB can be initialized to correspond to the memory applications of all tensors with the size of this memory address linked list block.

[0281] Based on the characteristics of relatively deterministic multi-round iterative training of deep learning tasks, the above technical solution proposes a heuristic memory reuse scheme combining two stages: the first stage is the first K mini-batches, which can also be called the policy learning stage. In this stage, according to the application and release requests of all tensors during each round of iterative training, an optimal memory reuse allocation policy is obtained by using a heuristic memory allocation search algorithm; the second stage is also called the policy application stage. In this stage, the optimal policy generated in the policy learning stage will be submitted to the memory pool allocator. Under the guidance of this allocation policy, the application and release requests of the next mini-batch for tensors are processed, thereby achieving the purpose of reducing memory and improving system performance. In existing deep learning frameworks, each tensor is usually only allocated to one cache block. Even if there is still enough remaining space in this cache block, other tensors cannot be allocated to this cache block. Only when this tensor is released can the cache block where it is located be used by other tensors, which directly leads to memory waste. And the solution of the present disclosure further improves the memory reuse efficiency and reduces the memory usage by placing a new tensor into a cache block that has been partially occupied by some tensors but still has enough space during allocation. Whenever a new tensor needs to apply for a cache block, the memory pool allocator starts searching from the memory address linked list block with a larger memory size than the memory size applied for by this tensor. If the free capacity of any cache block is greater than the memory size applied for by the current tensor, this tensor can be allocated to this cache block. If there is no cache block that meets the conditions, a new cache block will be created.

[0282] Figure 12 The flowchart showing a memory management method according to another embodiment of the present invention is as Figure 12 shown. The memory management method includes the following steps S1201 - S1203:

[0283] In step S1201, tensor application and release requests in the training data are obtained, where the tensor application and release requests carry tensor memory statistics information, and the tensor memory statistics information includes tensor size, tensor application time, and tensor release time;

[0284] In step S1202, the tensor memory statistics information is traversed, and the tensor memory statistics information is inserted into a cache block of a target memory address linked list block corresponding to the tensor size to obtain memory allocation information;

[0285] In step S1203, the memory pool allocator is initialized according to the memory allocation information, and tensor memory allocation is performed based on the initialized memory pool allocator.

[0286] In this embodiment, a memory management method is proposed. This method obtains memory allocation information based on representative training data, and then performs subsequent tensor memory allocation according to the memory allocation information. This technical solution is applicable not only to the allocation and management of small chunks of memory, but also to the allocation and management of large chunks of memory. It can not only perform memory allocation with lower system overhead, but also effectively improve the utilization rate of memory.

[0287] In an embodiment of the present invention, step S1202, that is, the step of traversing the tensor memory statistics information and inserting the tensor memory statistics information into a cache block of a target memory address linked list block corresponding to the tensor size to obtain memory allocation information, can be implemented as:

[0288] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes target tensor size, target tensor application time, and target tensor release time;

[0289] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistics information is stored in the cache block;

[0290] Determine whether there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information;

[0291] When there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not generate a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0292] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order;

[0293] In response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0294] In order to find memory allocation information with a suitable size for the tensor memory statistics information, while maximizing the reuse of memory, improving the utilization rate of memory, and reducing the waste of memory resources, in this embodiment, a traversal and trial method is adopted to obtain memory allocation information based on the tensor memory statistics information.

[0295] In an embodiment of the present invention, the step of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size can be implemented as:

[0296] Determine the memory address interval of the memory address linked list block;

[0297] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0298] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0299] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0300] In order to make full use of the storage space of the memory address linked list block while meeting the target tensor size requirement, in this embodiment, the target memory address linked list block corresponding to the target tensor size is determined based on the memory address interval of the memory address linked list block and the target tensor size.

[0301] Considering that some variable sparse tensors may cause the cache blocks determined according to the memory allocation information to fail to meet their memory requirements, in order to perform reasonable memory allocation for these special tensors, in this embodiment, if the memory allocation for the above-mentioned preset tensors fails, that is, a tensor memory allocation failure event is detected, the tensors with memory allocation failures are reallocated according to the preset memory reallocation rules. That is, in an embodiment of the present invention, the method further includes the step of reallocating the tensors with memory reallocation failures according to the preset memory allocation rules in response to detecting a tensor memory allocation failure event. Among them, the preset memory reallocation rules can be set according to the actual application needs and the data characteristics of the tensors, and the present disclosure does not make specific limitations on it.

[0302] In an embodiment of the present invention, the step of reallocating the tensors with memory allocation failures according to the preset memory allocation rules in response to detecting a tensor memory allocation failure event may include the following steps:

[0303] In response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the tensor with memory allocation failure according to the size of the tensor with memory allocation failure;

[0304] Determine whether there is a free cache block in the candidate memory address linked list block;

[0305] When there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the tensor with memory allocation failure, and fill the first header information for the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located;

[0306] When there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the tensor with memory allocation failure, and fill the second header information for the obtained memory, where the second header information stores the operating system memory allocation information.

[0307] In this embodiment, after detecting a tensor memory allocation failure event, first, a candidate memory address linked list block corresponding to the tensor size of the memory allocation failure is determined according to the tensor size of the memory allocation failure. The method for determining the candidate memory address linked list block can refer to the method for determining the target memory address linked list block above, which will not be elaborated herein. Then, it is determined whether there is a free cache block in the candidate memory address linked list block. When there is a free cache block in the candidate memory address linked list block, the free cache block is allocated to the tensor of the memory allocation failure, and first header information is filled in the free cache block. The first header information stores the memory address linked list block information where the free cache block is located, which is used to identify the storage and memory allocation location of the tensor of the memory allocation failure and provide a basis for subsequent memory release. When there is no free cache block in the candidate memory address linked list block, memory allocation is directly applied to the operating system based on the tensor size of the memory allocation failure, and second header information is filled in the obtained memory. The second header information stores the operating system memory allocation information, which is used to identify the storage and memory allocation location of the tensor of the memory allocation failure and provide a basis for subsequent memory release.

[0308] To improve the utilization rate of memory, after detecting a tensor memory release command, the occupied memory can be released according to a preset memory release rule. That is, in an embodiment of the present invention, the method further includes the step of releasing the memory according to a preset memory release rule in response to detecting a tensor memory release command. The preset memory release rule can be set according to the needs of actual applications and the data characteristics of tensors, and the present disclosure does not make specific limitations on it.

[0309] In an embodiment of the present invention, the step of releasing the memory according to a preset memory release rule in response to detecting a tensor memory release command may include the following steps:

[0310] In response to detecting a tensor memory release command, obtain the header information corresponding to the tensor memory release command;

[0311] Release the memory according to the header information.

[0312] As mentioned above, the header information is used to identify the storage and memory allocation location of the memory allocation failure tensor. Therefore, after detecting the tensor memory release command, the header information corresponding to the tensor memory release command can be obtained according to the pointer information of the released tensor, and then the memory can be released according to the header information, so that the memory can return to the idle state in the shortest time to store other information in a timely manner. For example, if the header information indicates that the tensor is stored in a cache block of a certain memory address linked list block, the released memory information is returned to the cache block of the memory address linked list block, otherwise it can be directly returned to the operating system.

[0313] Figure 13 The flowchart showing the memory management method according to another embodiment of the present invention is as Figure 13 shown, and the memory management method includes the following steps S1301 - S1303:

[0314] In step S1301, tensor memory statistical information is obtained, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0315] In step S1302, the tensor memory statistical information is traversed, and the tensor memory statistical information is inserted into the cache block of the target memory address linked list block corresponding to the tensor size;

[0316] In step S1303, in response to the end of the traversal of the tensor memory statistical information, the current memory address linked list block information is determined as the memory allocation information.

[0317] In this embodiment, a memory management method is proposed. This method traverses the representative tensor memory statistical information and inserts the tensor memory statistical information into the cache block of the target memory address linked list block corresponding to the tensor size to obtain the memory allocation information that can be used for tensor memory allocation subsequently. This technical solution is not only applicable to the allocation and management of small - sized memory, but also applicable to the allocation and management of large - sized memory. It can not only perform memory allocation with low system overhead, but also effectively improve the utilization rate of memory.

[0318] In an embodiment of the present invention, step S1302, that is, the step of traversing the tensor memory statistical information and inserting the tensor memory statistical information into the cache block of the target memory address linked list block corresponding to the tensor size, can be implemented as:

[0319] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes target tensor size, target tensor application time, and target tensor release time;

[0320] Determine a target memory address linked list block corresponding to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks;

[0321] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0322] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order;

[0323] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order.

[0324] In order to find memory allocation information with a suitable size for the tensor memory statistical information, and at the same time maximize the reuse of memory, improve the utilization rate of memory, and reduce the waste of memory resources, in this embodiment, a traversal and trial method is adopted to obtain memory allocation information based on the tensor memory statistical information.

[0325] In an embodiment of the present invention, the step of determining a target memory address linked list block corresponding to the target tensor size can be implemented as:

[0326] Determine the memory address interval of the memory address linked list block;

[0327] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0328] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0329] Determine the memory address linked list block corresponding to the size of the target memory address as the target memory address linked list block.

[0330] In order to make full use of the storage space of the memory address linked list block while meeting the requirements of the target tensor size, in this embodiment, the target memory address linked list block corresponding to the target tensor size is determined based on the memory address interval of the memory address linked list block and the target tensor size.

[0331] Considering that some variable sparse tensors may cause the cache blocks determined according to the memory allocation information to fail to meet their memory requirements, in order to also perform reasonable memory allocation for these special tensors, in this embodiment, if the memory allocation for the above preset tensor fails, that is, a tensor memory allocation failure event is detected, then the tensor with memory allocation failure is reallocated according to the preset memory reallocation rule. That is, in an embodiment of the present invention, the method further includes the step of reallocating the tensor with memory reallocation failure according to the preset memory allocation rule in response to detecting a tensor memory allocation failure event. Among them, the preset memory reallocation rule can be set according to the actual application needs and the data characteristics of the tensor, and the present disclosure does not make specific limitations on it.

[0332] In an embodiment of the present invention, the step of reallocating the tensor with memory allocation failure according to the preset memory allocation rule in response to detecting a tensor memory allocation failure event may include the following steps:

[0333] In response to detecting a tensor memory allocation failure event, determine the candidate memory address linked list block corresponding to the size of the tensor with memory allocation failure;

[0334] Determine whether there is an idle cache block in the candidate memory address linked list block;

[0335] When there is an idle cache block in the candidate memory address linked list block, allocate the idle cache block to the tensor with memory allocation failure, and fill the first header information for the idle cache block, where the first header information stores the information of the memory address linked list block where the idle cache block is located;

[0336] When there is no idle cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the tensor with memory allocation failure, and fill the second header information for the allocated memory, where the second header information stores the operating system memory allocation information.

[0337] In this embodiment, after detecting a tensor memory allocation failure event, first, a candidate memory address linked list block corresponding to the tensor size of the memory allocation failure is determined according to the tensor size of the memory allocation failure. The method for determining the candidate memory address linked list block can refer to the method for determining the target memory address linked list block above, which will not be elaborated herein; then, it is determined whether there is a free cache block in the candidate memory address linked list block; when there is a free cache block in the candidate memory address linked list block, the free cache block is allocated to the tensor of the memory allocation failure, and first header information is filled in the free cache block, where the first header information stores the memory address linked list block information where the free cache block is located, is used to identify the storage and memory allocation location of the tensor of the memory allocation failure, and provides a basis for subsequent memory release; when there is no free cache block in the candidate memory address linked list block, memory allocation is directly applied to the operating system based on the tensor size of the memory allocation failure, and second header information is filled in the obtained memory, where the second header information stores the operating system memory allocation information, is used to identify the storage and memory allocation location of the tensor of the memory allocation failure, and provides a basis for subsequent memory release.

[0338] Figures 12 - 13 The technical features in the embodiments shown are the same as or similar to those in the above Figures 1 - 11 embodiments shown. For the explanation and description of the technical features, reference can be made to the explanation and description of the above Figures 1 - 11 embodiments shown, which will not be elaborated herein.

[0339] The following is an embodiment of the device of the present invention, which can be used to execute the method embodiment of the present invention.

[0340] Figure 14 The structural block diagram of a memory management device according to an embodiment of the present invention is shown. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 14 shown, the memory management device includes:

[0341] A first acquisition module 1401, configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0342] A training module 1402, configured to train memory allocation information according to the tensor memory statistical information;

[0343] A first allocation module 1403, configured to perform tensor memory allocation according to the memory allocation information.

[0344] As mentioned above, with the development of data technology, deep learning has been widely applied in various fields, such as images, speech, advertising, search, recommendation, natural language processing, and so on. Researchers have found that memory management in deep learning frameworks is a key factor restricting the performance of deep learning. Currently, the optimization research on memory management mainly focuses on GPU video memory. However, in fact, the optimization of CPU cluster memory management is also crucial. This is because model training in CPU clusters has the characteristics of large scale and long training time, and a large amount of CPU and memory resources are required during the training process. Researchers found that a large amount of CPU resources were wasted on page faults during the performance evaluation process, which in turn led to memory waste.

[0345] In the prior art, memory management solutions for CPUs include the general Malloc library memory management solution, the combined memory management solution of BFCAllocator and Malloc library, and so on. Although these memory management solutions can improve the performance of network applications to a certain extent, they are not suitable for deep learning tasks. This is because deep learning tasks usually require frequent memory allocation and release. Especially for memory blocks of MB level, the total amount of operations may reach dozens of GB or even hundreds of GB. General memory management solutions generally only manage relatively small memory blocks, such as those less than 32KB. They do not have a cache design for large memory blocks and only allocate them in the operating system as needed. In this case, the page fault overhead caused by allocation in a frequently allocated scenario will seriously affect the system performance. If large memory blocks are cached, it may take a relatively long time before they can be reused, which will in turn cause relatively large resource waste. Especially in the cloud environment, the occupation of a large amount of memory means that the resources available to other applications running on the same machine will be significantly reduced, thus bringing mutual influence between applications. Therefore, there is an urgent need for a memory management method suitable for large memory blocks, which can not only allocate memory with low system overhead but also effectively improve the utilization rate of memory.

[0346] Considering the above problems, in this embodiment, a memory management device is proposed. This device trains memory allocation information based on representative training data, and then performs subsequent tensor memory allocation according to the memory allocation information. This technical solution is not only applicable to the allocation and management of small memory blocks but also to the allocation and management of large memory blocks. It can not only allocate memory with low system overhead but also effectively improve the utilization rate of memory.

[0347] Considering that although the cycle sizes and lifetimes of the memory blocks required in a mini-batch of deep learning tasks are different, the mini-batch data has cycle predictability, which makes the relative lifetimes across mini-batches relatively fixed. Herein, the mini-batch data refers to the several small training data sets obtained by dividing the entire training data set during the execution of the gradient descent algorithm in the deep learning scenario. Only one small training data set is trained each time. Since the gradient direction of the small training data set is not very different from that of the entire training data set, the correctness of the training can be ensured, and at the same time, the problem of huge computational complexity caused by all data sets participating in the training at one time can be avoided.

[0348] Based on the above data characteristics of deep learning tasks, in an embodiment of the present invention, memory allocation information is trained based on training data. Since the training data has a certain representativeness and can represent the characteristics of other data, the memory allocation information can be considered to be also applicable to other data. Subsequently, tensor memory allocation processing can be performed on other data according to the trained memory allocation information.

[0349] That is, in an embodiment of the present invention, the training data refers to the data used to train the memory allocation information that can be used for other data processing. Specifically, the training data can be the first K mini-batches of data in a certain deep learning task. Among them, the actual requests for tensor application and release generated in the training data participate in the training of the memory allocation information. Herein, the tensor refers to an array of arbitrary number of dimensions composed of a set of primitive values and is the data unit for storing data in TensorFlow. The tensor application and release requests carry tensor memory statistical information associated with memory allocation and decisive for the memory allocation strategy. The tensor memory statistical information at least includes information such as tensor size, tensor application time, and tensor release time. The tensor size is used to determine the size of the required memory space, and the tensor application time and tensor release time are used to determine the lifetime of the tensor memory statistical information, which can further be used as a basis for judging whether the memory space corresponding to the tensor size is reusable. That is, if there is an intersection in the request lifetimes of a certain memory space, it means that the memory space is not reusable for the requests with intersections. On the contrary, if there is no intersection in the request lifetimes of a certain memory space, it means that the memory space is reusable for the relevant requests.

[0350] In an embodiment of the present invention, the tensor memory statistical information may include one or more tensor memory statistical information. Subsequently, when training the memory allocation information, training will be performed based on all the tensor memory statistical information.

[0351] In an embodiment of the present invention, as Figure 15 shown, the training module 1402 includes:

[0352] A first determination sub-module 1501, configured to determine a preset traversal order;

[0353] A training sub-module 1502, configured to traverse the tensor memory statistics according to the preset traversal order, and train memory allocation information based on the tensor memory statistics.

[0354] As mentioned above, when training memory allocation information based on the tensor memory statistics, training is performed based on all the tensor memory statistics, that is, it is necessary to traverse all the information in the tensor memory statistics so that each tensor memory statistics contributes to the generation of the memory allocation information. Specifically, first, determine a preset traversal order; then, traverse the tensor memory statistics according to the preset traversal order, and train memory allocation information based on all the tensor memory statistics.

[0355] In an embodiment of the present invention, the preset traversal order can be set according to the needs of actual applications and the data characteristics of the tensor memory statistics, and the present disclosure does not make specific limitations thereon. For example, the preset traversal order can be in the order from largest to smallest tensor size, and of course, it can also be in the order from smallest to largest tensor size or in a randomly selected order according to tensor size, etc.

[0356] In an embodiment of the present invention, the training sub-module 1502 can be configured to:

[0357] Determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0358] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistics are stored in the cache blocks;

[0359] Determine whether there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information;

[0360] When there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not generate a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0361] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order;

[0362] In response to the end of traversing the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0363] In order to find memory allocation information of appropriate size for the tensor memory statistics information, while maximizing the reuse of memory, improving the utilization rate of memory, and reducing the waste of memory resources, in this embodiment, the memory allocation information is obtained by training based on the tensor memory statistics information in a way of traversing and probing.

[0364] Specifically:

[0365] First, determine the target tensor information that currently requires memory allocation according to the preset traversal order, that is, traverse the tensor information in the order of decreasing tensor size. Among them, the target tensor information at least includes information such as the target tensor size, the target tensor application time, and the target tensor release time.

[0366] Then, determine the target memory address linked list block corresponding to the target tensor size according to the target tensor size. Among them, one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block. Specifically, the size of the cache block should be greater than or equal to the target tensor size to sufficiently meet the memory storage requirements of the target tensor. Among them, the cache block may be an empty cache block or may already store one or more historical tensor memory statistics information. Subsequently, it is determined whether the target tensor information can be inserted into the cache block according to the comparison of the lifetimes between the target tensor information and the historical tensor memory statistics information.

[0367] Then, based on the lifetime comparison mentioned above, it is determined whether there is a cache block in the target memory address linked list block that does not have a preset conflict with the target tensor information. In an embodiment of the present invention, the preset conflict refers to a lifetime conflict. As mentioned above, if there is an intersection between the lifetime of the target tensor information and the lifetime of the historical tensor memory statistical information already stored in the cache block, it indicates that there is a lifetime conflict between the target tensor information and the cache block, and the memory space stored in this cache block is not reusable for the target tensor information. On the contrary, if there is no intersection between the lifetime of the target tensor information and the lifetime of the historical tensor memory statistical information already stored in the cache block, it indicates that there is no lifetime conflict between the target tensor information and the cache block, and the memory space stored in this cache block is reusable for the target tensor information.

[0368] Therefore, if there is a cache block in the target memory address linked list block that does not have a preset conflict with the target tensor information, it means that the memory space stored in this cache block is reusable for the target tensor information. The target tensor information can be inserted into the cache block that does not have a preset conflict with the target tensor information, and the memory allocation relationship for the current target tensor information is ended. The next target tensor information can be determined according to the preset traversal order to continue the memory allocation.

[0369] If there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, it means that the memory space stored in this cache block is not reusable for the target tensor information and cannot receive the memory allocation request corresponding to the target tensor information. At this time, it is necessary to find a cache block in other memory address linked list blocks that can receive the memory allocation request corresponding to the target tensor information. To improve the search efficiency, in an embodiment of the present invention, within the range of other memory address linked list blocks larger than the target memory address linked list block, traverse the memory address linked list blocks in ascending order of cache block size to find a cache block that can receive the memory allocation request corresponding to the target tensor information. Further, similar to the above description, if there is a cache block in the memory address linked list blocks within the search range that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and create a shared pointer in the target memory address linked list block to point to this cache block, return the traversal process of the tensor information, and determine the next target tensor information for processing according to the preset traversal order. However, if there is no cache block in the memory address linked list blocks that does not cause a preset conflict with the target tensor information, only the target memory address linked list block can be returned, create a new cache block in the target memory address linked list block to receive the memory allocation request corresponding to the target tensor information, insert the target tensor information into the new cache block, and return the traversal process of the tensor information, and determine the next target tensor information for processing according to the preset traversal order.

[0370] Finally, after the traversal of the tensor memory statistics information is completed, the obtained current memory address linked list block information is determined as the memory allocation information, and the memory allocation information can be used for memory allocation for other data.

[0371] In this way, not only can a memory allocation information with a suitable size be found for the tensor memory statistics information, but also the reuse of memory can be maximally realized, the utilization rate of memory can be improved, and the waste of memory resources can be reduced.

[0372] In an embodiment of the present invention, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size can be implemented as:

[0373] Determine the memory address interval of the memory address linked list block;

[0374] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0375] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0376] Determine the memory address linked list block corresponding to the size of the target memory address as the target memory address linked list block.

[0377] In order to make full use of the storage space of the memory address linked list block while meeting the requirements of the target tensor size, in this embodiment, the target memory address linked list block corresponding to the target tensor size is determined based on the memory address interval of the memory address linked list block and the target tensor size. Specifically, first determine the memory address interval of the memory address linked list block, where the memory address interval can be set according to the needs of actual applications, and the present disclosure does not make specific limitations on it; then round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; then multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; finally, determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0378] For example, if the memory address interval is set to 4KB and the target tensor size is 33KB, 33KB / 4KB = 8.25, then the required number of memory address intervals is 9, and the target memory address size is 4KB × 9 = 36KB.

[0379] In an embodiment of the present invention, as Figure 16 shown, the first allocation module 1403 includes:

[0380] An initialization sub-module 1601, configured to initialize the memory pool allocator according to the memory allocation information;

[0381] A first allocation sub-module 1602, configured to perform tensor memory allocation based on the initialized memory pool allocator.

[0382] After obtaining the memory allocation information through training, tensor memory allocation processing can be performed on other data according to the memory allocation information. Specifically, the initialization sub-module 1101 initializes the memory pool allocator according to the memory allocation information; the first allocation sub-module 1102 performs tensor memory allocation based on the initialized memory pool allocator. For example, if there are two cache blocks in the memory address linked list block with a cache block size of 36KB, then 2 caches with a size of 36KB can be initialized in the cache of the memory pool allocator for memory applications corresponding to all tensors of the size of this memory address linked list block.

[0383] In an embodiment of the present invention, the device further includes a part that, in response to detecting a tensor memory allocation failure event, re-allocates the memory allocation failure tensor according to a preset memory allocation rule, that is, as Figure 17As shown in the figure, the memory management device includes:

[0384] A first acquisition module 1701, configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0385] A training module 1702, configured to train memory allocation information according to the tensor memory statistical information;

[0386] A first allocation module 1703, configured to perform tensor memory allocation according to the memory allocation information;

[0387] A first reallocation module 1704, configured to, in response to detecting a tensor memory allocation failure event, reallocate tensors that fail in memory reallocation according to a preset memory allocation rule.

[0388] Considering that some variable sparse tensors may cause the cache blocks determined according to the memory allocation information to be unable to meet their memory requirements, in order to perform reasonable memory allocation for these special tensors as well, in this embodiment, if the memory allocation for the above-mentioned preset tensors fails, that is, a tensor memory allocation failure event is detected, the first reallocation module 1704 reallocates the tensors that fail in memory allocation according to a preset memory reallocation rule.

[0389] In an embodiment of the present invention, the preset memory reallocation rule can be set according to the actual application needs and the data characteristics of the tensors, and the present disclosure does not make specific limitations on it.

[0390] In an embodiment of the present invention, as Figure 18 shown, the first reallocation module 1704 includes:

[0391] A second determination sub-module 1801, configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the tensor that fails in memory allocation according to the size of the tensor that fails in memory allocation;

[0392] A third determination sub-module 1802, configured to determine whether there is an idle cache block in the candidate memory address linked list block;

[0393] A second allocation sub-module 1803, configured to, when there is an idle cache block in the candidate memory address linked list block, allocate the idle cache block to the tensor that fails in memory allocation, and fill the first header information in the idle cache block, where the first header information stores the information of the memory address linked list block where the idle cache block is located;

[0394] The third allocation sub-module 1804 is configured to apply for memory allocation from the operating system based on the memory allocation failure tensor size when there is no free cache block in the candidate memory address linked list block, and fill the second header information for the allocated memory, where the operating system memory allocation information is stored in the second header information.

[0395] In this embodiment, after detecting the tensor memory allocation failure event, the second determination sub-module 1801 determines a candidate memory address linked list block corresponding to the memory allocation failure tensor size according to the memory allocation failure tensor size. The determination method of the candidate memory address linked list block may refer to the determination method of the target memory address linked list block above, which will not be elaborated herein. The third determination sub-module 1802 determines whether there is a free cache block in the candidate memory address linked list block. When there is a free cache block in the candidate memory address linked list block, the second allocation sub-module 1803 allocates the free cache block to the memory allocation failure tensor, and fills the first header information for the free cache block, where the memory address linked list block information where the free cache block is located is stored in the first header information, which is used to identify the storage and memory allocation location of the memory allocation failure tensor and provide a basis for subsequent memory release. When there is no free cache block in the candidate memory address linked list block, the third allocation sub-module 1804 directly applies for memory allocation from the operating system based on the memory allocation failure tensor size, and fills the second header information for the allocated memory, where the operating system memory allocation information is stored in the second header information, which is used to identify the storage and memory allocation location of the memory allocation failure tensor and provide a basis for subsequent memory release.

[0396] In an embodiment of the present invention, the device further includes a part for releasing memory according to a preset memory release rule in response to detecting a tensor memory release command, that is, as Figure 19 shown, the memory management device includes:

[0397] The first acquisition module 1901 is configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0398] The training module 1902 is configured to train memory allocation information according to the tensor memory statistical information;

[0399] The first allocation module 1903 is configured to perform tensor memory allocation according to the memory allocation information;

[0400] The first redistribution module 1904 is configured to, in response to detecting a tensor memory allocation failure event, redistribute the memory redistribution failure tensor according to a preset memory allocation rule;

[0401] The first release module 1905 is configured to, in response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

[0402] In this embodiment, in order to improve the utilization rate of the memory, after detecting the tensor memory release command, the first release module 1905 can release the occupied memory according to the preset memory release rule.

[0403] In an embodiment of the present invention, the preset memory release rule can be set according to the needs of the actual application and the data characteristics of the tensor, and the present disclosure does not make specific limitations thereto.

[0404] In an embodiment of the present invention, as Figure 20 shown, the first release module 1905 includes:

[0405] The first acquisition sub-module 2001 is configured to, in response to detecting a tensor memory release command, acquire the header information corresponding to the tensor memory release command;

[0406] The first release sub-module 2002 is configured to release the memory according to the header information.

[0407] As mentioned above, the header information is used to identify the storage and memory allocation location of the memory allocation failure tensor. Therefore, after detecting the tensor memory release command, the first acquisition sub-module 2001 can acquire the header information corresponding to the tensor memory release command according to the pointer information of the tensor release, and the first release sub-module 2002 then releases the memory according to the header information, so that the memory can return to the idle state in the shortest time to store other information in a timely manner. For example, if the header information indicates that the tensor is stored in a cache block of a certain memory address linked list block, the released memory information is returned to the cache block of the memory address linked list block, otherwise it can be directly returned to the operating system.

[0408] Figure 21 The structural block diagram of a memory management device according to another embodiment of the present invention is shown. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 21 shown, the memory management device includes:

[0409] A second acquisition module 2101, configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0410] A first insertion module 2102, configured to traverse the tensor memory statistical information, insert the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size, and obtain memory allocation information;

[0411] A second allocation module 2103, configured to initialize a memory pool allocator according to the memory allocation information, and perform tensor memory allocation based on the initialized memory pool allocator.

[0412] In this embodiment, a memory management device is proposed. The device obtains memory allocation information based on representative training data, and then performs subsequent tensor memory allocation according to the memory allocation information. This technical solution is applicable not only to the allocation and management of small chunks of memory, but also to the allocation and management of large chunks of memory. It can not only perform memory allocation with low system overhead, but also effectively improve the utilization rate of memory.

[0413] In an embodiment of the present invention, the first insertion module 2102 may be configured to:

[0414] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes target tensor size, target tensor application time, and target tensor release time;

[0415] Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks;

[0416] Determine whether there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information;

[0417] When there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not generate a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order;

[0418] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block in ascending order. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order;

[0419] In response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

[0420] In order to find memory allocation information of a suitable size for the tensor memory statistics information, and at the same time maximize the reuse of memory, improve the utilization rate of memory, and reduce the waste of memory resources, in this embodiment, the first insertion module 2102 obtains the memory allocation information based on the tensor memory statistics information in a way of traversing and probing.

[0421] In an embodiment of the present invention, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size may be configured as:

[0422] Determine the memory address interval of the memory address linked list block;

[0423] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0424] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0425] Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

[0426] In order to make full use of the storage space of the memory address linked list block while meeting the requirements of the target tensor size, in this embodiment, the target memory address linked list block corresponding to the target tensor size is determined based on the memory address interval of the memory address linked list block and the target tensor size.

[0427] Considering that the cache blocks determined according to the memory allocation information may not meet the memory requirements of some variable sparse tensors, in order to reasonably allocate memory for these special tensors, in this embodiment, if the memory allocation for the above-mentioned preset tensors fails, that is, a tensor memory allocation failure event is detected, the tensors with memory allocation failures are reallocated according to the preset memory reallocation rules. That is, in an embodiment of the present invention, the device further includes a second reallocation module that, in response to detecting a tensor memory allocation failure event, reallocates the tensors with failed memory reallocation according to the preset memory allocation rules. Among them, the preset memory reallocation rules can be set according to the actual application needs and the data characteristics of the tensors, and the present disclosure does not make specific limitations on it.

[0428] In an embodiment of the present invention, the second reallocation module may include:

[0429] A fourth determination sub-module, configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the tensor with memory allocation failure according to the size of the tensor with memory allocation failure;

[0430] A fifth determination sub-module, configured to determine whether there is a free cache block in the candidate memory address linked list block;

[0431] A fourth allocation sub-module, configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the tensor with memory allocation failure, and fill the first header information in the free cache block, where the first header information stores the information of the memory address linked list block where the free cache block is located;

[0432] A fifth allocation sub-module, configured to, when there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the tensor with memory allocation failure, and fill the second header information in the allocated memory, where the second header information stores the operating system memory allocation information.

[0433] In this embodiment, after detecting a tensor memory allocation failure event, the fourth determination sub-module determines a candidate memory address linked list block corresponding to the memory allocation failure tensor size according to the memory allocation failure tensor size. The determination method of the candidate memory address linked list block can refer to the determination method of the target memory address linked list block above, and the present disclosure will not elaborate here; the fifth determination sub-module determines whether there is a free cache block in the candidate memory address linked list block; when there is a free cache block in the candidate memory address linked list block, the fourth allocation sub-module allocates the free cache block to the memory allocation failure tensor and fills the first header information into the free cache block. The first header information stores the memory address linked list block information where the free cache block is located, which is used to identify the storage and memory allocation location of the memory allocation failure tensor and provide a basis for subsequent memory release; when there is no free cache block in the candidate memory address linked list block, the fifth allocation sub-module directly applies to the operating system for memory allocation based on the memory allocation failure tensor size and fills the second header information into the allocated memory. The second header information stores the operating system memory allocation information, which is used to identify the storage and memory allocation location of the memory allocation failure tensor and provide a basis for subsequent memory release.

[0434] To improve the utilization rate of memory, after detecting a tensor memory release command, the occupied memory can be released according to a preset memory release rule. That is, in an embodiment of the present invention, the device further includes a second release module that responds to detecting a tensor memory release command and releases the memory according to a preset memory release rule. The preset memory release rule can be set according to the actual application needs and the data characteristics of the tensor, and the present disclosure does not specifically limit it.

[0435] In an embodiment of the present invention, the second release module may include:

[0436] A second acquisition sub-module configured to respond to detecting a tensor memory release command and acquire the header information corresponding to the tensor memory release command;

[0437] A second release sub-module configured to release the memory according to the header information.

[0438] As mentioned above, the header information is used to identify the storage and memory allocation location of the memory allocation failure tensor. Therefore, after detecting the tensor memory release command, the second acquisition sub-module acquires the header information corresponding to the tensor memory release command according to the pointer information of the tensor release, and then the second release sub-module releases the memory according to the header information, so that the memory can return to the idle state in the shortest time to store other information in a timely manner. For example, if the header information indicates that the tensor is stored in a cache block of a certain memory address linked list block, the released memory information is returned to the cache block of the memory address linked list block, otherwise it can be directly returned to the operating system.

[0439] Figure 22 FIG. shows a structural block diagram of a memory management device according to another embodiment of the present invention. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 22 shown, the memory management device includes:

[0440] A third acquisition module 2201, configured to acquire tensor memory statistical information, where the tensor memory statistical information includes tensor size, tensor application time, and tensor release time;

[0441] A second insertion module 2202, configured to traverse the tensor memory statistical information and insert the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size;

[0442] A determination module 2203, configured to, in response to the end of the traversal of the tensor memory statistical information, determine the current memory address linked list block information as the memory allocation information.

[0443] In this embodiment, a memory management device is proposed. The device traverses representative tensor memory statistical information and inserts the tensor memory statistical information into a cache block of a target memory address linked list block corresponding to the tensor size to obtain subsequent memory allocation information that can be used for tensor memory allocation. This technical solution is applicable not only to the allocation and management of small memory blocks, but also to the allocation and management of large memory blocks. It can not only perform memory allocation with low system overhead, but also effectively improve the utilization rate of memory.

[0444] In an embodiment of the present invention, the second insertion module 2202 may be configured to:

[0445] Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time;

[0446] Determine a target memory address linked list block corresponding to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks;

[0447] Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information;

[0448] When there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order;

[0449] When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks whose sizes are larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order.

[0450] In order to find memory allocation information with a suitable size for the tensor memory statistical information, and at the same time maximize the reuse of memory, improve the utilization rate of memory, and reduce the waste of memory resources, in this embodiment, the second insertion module 2202 obtains memory allocation information based on the tensor memory statistical information in a way of traversing and probing.

[0451] In an embodiment of the present invention, the part of determining the target memory address linked list block corresponding to the target tensor size may be configured as:

[0452] Determine the memory address interval of the memory address linked list block;

[0453] Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals;

[0454] Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size;

[0455] Determine the memory address linked list block corresponding to the size of the target memory address as the target memory address linked list block.

[0456] In order to make full use of the storage space of the memory address linked list block while meeting the requirements of the target tensor size, in this embodiment, the target memory address linked list block corresponding to the target tensor size is determined based on the memory address interval of the memory address linked list block and the target tensor size.

[0457] Considering that some variable sparse tensors may cause the cache blocks determined according to the memory allocation information to fail to meet their memory requirements, in order to perform reasonable memory allocation for these special tensors, in this embodiment, if the memory allocation for the above-mentioned preset tensors fails, that is, a tensor memory allocation failure event is detected, then the tensors with memory allocation failures are reallocated according to the preset memory reallocation rules. That is, in an embodiment of the present invention, the device further includes a third reallocation module that, in response to detecting a tensor memory allocation failure event, reallocates the tensors with failed memory reallocation according to the preset memory allocation rules. Among them, the preset memory reallocation rules can be set according to the actual application needs and the data characteristics of the tensors, and the present disclosure does not make specific limitations on it.

[0458] In an embodiment of the present invention, the third reallocation module may include:

[0459] A sixth determination sub-module, configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the tensor with memory allocation failure according to the size of the tensor with memory allocation failure;

[0460] A seventh determination sub-module, configured to determine whether there is a free cache block in the candidate memory address linked list block;

[0461] A sixth allocation sub-module, configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the tensor with memory allocation failure, and fill the first header information in the free cache block, where the first header information stores the information of the memory address linked list block where the free cache block is located;

[0462] A seventh allocation sub-module, configured to, when there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the tensor with memory allocation failure, and fill the second header information in the applied memory, where the second header information stores the operating system memory allocation information.

[0463] In this embodiment, after detecting a tensor memory allocation failure event, the sixth determination sub-module determines a candidate memory address linked list block corresponding to the memory allocation failure tensor size according to the memory allocation failure tensor size. The method for determining the candidate memory address linked list block can refer to the method for determining the target memory address linked list block above, which will not be elaborated herein; the seventh determination sub-module determines whether there is an idle cache block in the candidate memory address linked list block; when there is an idle cache block in the candidate memory address linked list block, the sixth allocation sub-module allocates the idle cache block to the memory allocation failure tensor and fills the first header information into the idle cache block, where the first header information stores the memory address linked list block information where the idle cache block is located, is used to identify the storage and memory allocation location of the memory allocation failure tensor, and provides a basis for subsequent memory release; when there is no idle cache block in the candidate memory address linked list block, the seventh allocation sub-module directly applies to the operating system for memory allocation based on the memory allocation failure tensor size and fills the second header information into the allocated memory, where the second header information stores the operating system memory allocation information, is used to identify the storage and memory allocation location of the memory allocation failure tensor, and provides a basis for subsequent memory release.

[0464] Figures 21 - 22 The technical features in the Figures 12 - 20 shown embodiment are the same as or similar to the technical features in the above Figures 12 - 20 shown embodiment. For the explanation and description of the technical features, reference can be made to the explanation and description of the

[0465] shown embodiment above, which will not be elaborated herein. Figure 23 The structural block diagram of an electronic device according to an embodiment of the present invention is shown, as Figure 23 shown, the electronic device 2300 includes a memory 2301 and a processor 2302; wherein,

[0466] The memory 2301 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor 2302 to implement any of the above method steps.

[0467] Figure 24 The structural schematic diagram of a computer system suitable for implementing the memory management method according to an embodiment of the present invention.

[0468] As Figure 24As shown, the computer system 2400 includes a processing unit 2401, which can perform various processes in the above-described embodiments according to a program stored in a read-only memory (ROM) 2402 or a program loaded from a storage section 2408 into a random access memory (RAM) 2403. In the RAM 2403, various programs and data required for the operation of the system 2400 are also stored. The processing unit 2401, the ROM 2402, and the RAM 2403 are connected to each other via a bus 2404. An input / output (I / O) interface 2405 is also connected to the bus 2404.

[0469] The following components are connected to the I / O interface 2405: an input section 2406 including a keyboard, a mouse, etc.; an output section 2407 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 2408 including a hard disk, etc.; and a communication section 2409 including a network interface card such as a LAN card, a modem, etc. The communication section 2409 performs communication processing via a network such as the Internet. A drive 2410 is also connected to the I / O interface 2405 as needed. A removable medium 2411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 2410 as needed so that a computer program read from it can be installed into the storage section 2408 as needed. Among them, the processing unit 2401 can be implemented as a processing unit such as a CPU, a GPU, an FPAG, an NPU, etc.

[0470] In particular, according to an embodiment of the present invention, the method described above can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly contained on a computer-readable medium, and the computer program includes program code for executing the memory management method. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 2409, and / or installed from the removable medium 2411.

[0471] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0472] The units or modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor, and the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.

[0473] As another aspect, the embodiments of the present invention also provide a computer-readable storage medium, which can be the computer-readable storage medium included in the device in the above embodiments; or can be a computer-readable storage medium that exists alone and is not assembled into the device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the embodiments of the present invention.

[0474] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present invention.

Claims

1. A memory management method, characterized in that, it includes: Obtain tensor application and release requests in the training data, where the tensor application and release requests carry tensor memory statistical information, and the tensor memory statistical information includes tensor size, tensor application time, and tensor release time. The training data includes at least one mini-batch of data for a deep learning task; Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes target tensor size, target tensor application time, and target tensor release time; Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks; Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information; When there is such a cache block in the target memory address linked list block, insert the target tensor information into the cache block, and determine the next target tensor information for processing according to the preset traversal order; In response to the end of traversing the tensor memory statistical information, determine the current memory address linked list block information as the memory allocation information; Perform tensor memory allocation according to the memory allocation information.

2. The method according to claim 1, characterized in that, the method further includes: When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order.

3. The method according to claim 1, characterized in that, the step of determining a target memory address linked list block corresponding to the target tensor size according to the target tensor size is implemented as: Determine the memory address interval of the memory address linked list block; Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

4. The method according to any one of claims 1-3, characterized in that, the step of performing tensor memory allocation according to the memory allocation information includes: Initialize the memory pool allocator according to the memory allocation information; Perform tensor memory allocation based on the initialized memory pool allocator.

5. The method according to any one of claims 1-3, characterized in that, further comprising: In response to detecting a tensor memory allocation failure event, reallocate the memory reallocation failure tensor according to a preset memory allocation rule.

6. The method according to claim 5, characterized in that, The step of, in response to detecting a tensor memory allocation failure event, reallocating the memory allocation failure tensor according to a preset memory allocation rule, includes: In response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failure tensor according to the size of the memory allocation failure tensor; Determine whether there is a free cache block in the candidate memory address linked list block; When there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill a first header information for the free cache block, wherein the first header information stores the memory address linked list block information where the free cache block is located; When there is no free cache block in the candidate memory address linked list block, apply for memory allocation from the operating system based on the size of the memory allocation failure tensor, and fill a second header information for the obtained memory, wherein the second header information stores the operating system memory allocation information.

7. The method according to claim 6, characterized in that, further comprising: In response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

8. The method according to claim 7, characterized in that, The step of, in response to detecting a tensor memory release command, releasing the memory according to a preset memory release rule, includes: In response to detecting a tensor memory release command, obtain the header information corresponding to the tensor memory release command; Release the memory according to the header information.

9. A memory management method, characterized in that, comprising: Obtain tensor application and release requests in the training data, wherein the tensor application and release requests carry tensor memory statistical information, the tensor memory statistical information includes tensor size, tensor application time and tensor release time, and the training data includes at least one mini-batch data of a deep learning task; Determine a preset traversal order, and determine target tensor information according to the preset traversal order, wherein the target tensor information includes target tensor size, target tensor application time and target tensor release time; Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, wherein one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache blocks; Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information; When there is such a cache block in the target memory address linked list block, insert the target tensor information into the cache block, and determine the next target tensor information to be processed according to the preset traversal order; In response to the end of traversing the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information; Initialize the memory pool allocator according to the memory allocation information, and perform tensor memory allocation based on the initialized memory pool allocator.

10. The method according to claim 9, wherein, the method further includes: When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order; In response to the end of traversing the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

11. The method according to claim 9, wherein, the determining the target memory address linked list block corresponding to the target tensor size is implemented as: Determine the memory address interval of the memory address linked list block; Round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; Multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; Determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

12. The method according to any one of claims 9-11, wherein, it further includes: In response to detecting a tensor memory allocation failure event, re-allocate the memory re-allocation failed tensor according to a preset memory allocation rule.

13. The method according to claim 12, wherein, the responding to detecting a tensor memory allocation failure event and re-allocating the memory allocation failed tensor according to a preset memory allocation rule includes: In response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failed tensor according to the size of the memory allocation failed tensor; Determine whether there is an idle cache block in the candidate memory address linked list block; When there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failed tensor, and fill the first header information into the free cache block, where the first header information stores the information of the memory address linked list block where the free cache block is located; When there is no free cache block in the candidate memory address linked list block, apply for memory allocation from the operating system based on the size of the memory allocation failed tensor, and fill the second header information into the obtained memory, where the second header information stores the operating system memory allocation information.

14. The method according to claim 13, wherein, further comprising: In response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

15. The method according to claim 14, wherein, the step of in response to detecting a tensor memory release command, releasing the memory according to a preset memory release rule includes: In response to detecting a tensor memory release command, obtain the header information corresponding to the tensor memory release command; Release the memory according to the header information.

16. A memory management method, wherein, comprising: Obtain the tensor memory statistical information in the training data, where the tensor memory statistical information includes the tensor size, the tensor application time, and the tensor release time, and the training data includes at least one mini-batch data of a deep learning task; Determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes the target tensor size, the target tensor application time, and the target tensor release time; Determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache block; Determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information; When there is such a cache block in the target memory address linked list block, insert the target tensor information into the cache block, and determine the next target tensor information for processing according to the preset traversal order; In response to the end of the traversal of the tensor memory statistical information, determine the current memory address linked list block information as the memory allocation information.

17. The method according to claim 16, wherein, the method further comprises: When there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from smallest to largest. When there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information to be processed according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information to be processed according to the preset traversal order.

18. The method according to claim 16, wherein, the determining the target memory address linked list block corresponding to the target tensor size is implemented as: determine the memory address interval of the memory address linked list block; round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

19. The method according to any one of claims 16-18, wherein, further comprising: in response to detecting a tensor memory allocation failure event, reallocate the memory allocation failed tensor according to a preset memory allocation rule.

20. The method according to claim 19, wherein, the responding to detecting a tensor memory allocation failure event and reallocating the memory allocation failed tensor according to a preset memory allocation rule includes: in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failed tensor; determine whether there is an idle cache block in the candidate memory address linked list block; when there is an idle cache block in the candidate memory address linked list block, allocate the idle cache block to the memory allocation failed tensor, and fill the first header information in the idle cache block, wherein the first header information stores the memory address linked list block information where the idle cache block is located; when there is no idle cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the memory allocation failed tensor, and fill the second header information in the allocated memory, wherein the second header information stores the operating system memory allocation information.

21. A memory management device, wherein, comprising: A first acquisition module, configured to acquire tensor application and release requests in training data, where the tensor application and release requests carry tensor memory statistics, and the tensor memory statistics include tensor size, tensor application time, and tensor release time, and the training data includes at least one mini-batch of data for a deep learning task; A training module, configured to determine a preset traversal order, and determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time; determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistics are stored in the cache blocks; determine whether there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information; when there is such a cache block in the target memory address linked list block, insert the target tensor information into the cache block, and determine the next target tensor information for processing according to the preset traversal order; in response to the end of the traversal of the tensor memory statistics, determine the current memory address linked list block information as memory allocation information; A first allocation module, configured to perform tensor memory allocation according to the memory allocation information.

22. The apparatus according to claim 21, wherein, the training module is configured to: when there is no cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large, and when there is a cache block in the memory address linked list block that does not generate a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not generate a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order, and when there is no cache block in the memory address linked list block that does not generate a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order; in response to the end of the traversal of the tensor memory statistics, determine the current memory address linked list block information as the memory allocation information.

23. The apparatus according to claim 21, wherein, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is configured to: determine the memory address interval of the memory address linked list block; round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

24. The device according to any one of claims 21-23, characterized in that, the first allocation module includes: an initialization sub-module configured to initialize a memory pool allocator according to the memory allocation information; a first allocation sub-module configured to perform tensor memory allocation based on the initialized memory pool allocator.

25. The device according to any one of claims 21-23, characterized in that, it further includes: a first reallocation module configured to, in response to detecting a tensor memory allocation failure event, reallocate a memory reallocation failure tensor according to a preset memory allocation rule.

26. The device according to claim 25, characterized in that, the first reallocation module includes: a second determination sub-module configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failure tensor according to the size of the memory allocation failure tensor; a third determination sub-module configured to determine whether there is a free cache block in the candidate memory address linked list block; a second allocation sub-module configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor and fill a first header information into the free cache block, wherein the first header information stores information about the memory address linked list block where the free cache block is located; a third allocation sub-module configured to, when there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the memory allocation failure tensor and fill a second header information into the allocated memory, wherein the second header information stores operating system memory allocation information.

27. The device according to claim 26, characterized in that, it further includes: a first release module configured to, in response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

28. The device according to claim 27, characterized in that, the first release module includes: a first acquisition sub-module configured to, in response to detecting a tensor memory release command, acquire header information corresponding to the tensor memory release command; a first release sub-module configured to release the memory according to the header information.

29. A memory management device, characterized in that, it includes: a second acquisition module configured to acquire tensor application and release requests in training data, wherein the tensor application and release requests carry tensor memory statistical information, the tensor memory statistical information includes tensor size, tensor application time and tensor release time, and the training data includes at least one mini-batch data of a deep learning task; The first insertion module is configured to determine a preset traversal order, determine target tensor information according to the preset traversal order, where the target tensor information includes a target tensor size, a target tensor application time, and a target tensor release time; determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, where one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistics information is stored in the cache blocks; determine whether there is a cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information; when there is such a cache block in the target memory address linked list block, insert the target tensor information into the cache block, and determine the next target tensor information for processing according to the preset traversal order; in response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as memory allocation information; The second allocation module is configured to initialize a memory pool allocator according to the memory allocation information, and perform tensor memory allocation based on the initialized memory pool allocator.

30. The apparatus according to claim 29, wherein, the first insertion module is configured to: when there is no cache block in the target memory address linked list block that does not generate a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large, when there is a cache block in the memory address linked list block that does not generate a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not generate a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order, when there is no cache block in the memory address linked list block that does not generate a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order; in response to the end of the traversal of the tensor memory statistics information, determine the current memory address linked list block information as the memory allocation information.

31. The apparatus according to claim 29, wherein, the part of determining the target memory address linked list block corresponding to the target tensor size according to the target tensor size is configured to: determine the memory address interval of the memory address linked list block; round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; multiply the memory address interval by the required number of memory address intervals to obtain a target memory address size; determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

32. The apparatus according to any one of claims 29 - 31, wherein, further comprising: A second redistribution module, configured to, in response to detecting a tensor memory allocation failure event, redistribute a memory redistribution failure tensor according to a preset memory allocation rule.

33. The apparatus according to claim 32, wherein, the second redistribution module includes: a fourth determination sub-module, configured to, in response to detecting a tensor memory allocation failure event, determine a candidate memory address linked list block corresponding to the size of the memory allocation failure tensor according to the size of the memory allocation failure tensor; a fifth determination sub-module, configured to determine whether there is a free cache block in the candidate memory address linked list block; a fourth allocation sub-module, configured to, when there is a free cache block in the candidate memory address linked list block, allocate the free cache block to the memory allocation failure tensor, and fill a first header information into the free cache block, wherein the first header information stores information about the memory address linked list block where the free cache block is located; a fifth allocation sub-module, configured to, when there is no free cache block in the candidate memory address linked list block, apply to the operating system for memory allocation based on the size of the memory allocation failure tensor, and fill a second header information into the allocated memory, wherein the second header information stores operating system memory allocation information.

34. The apparatus according to claim 33, wherein, it further includes: a second release module, configured to, in response to detecting a tensor memory release command, release the memory according to a preset memory release rule.

35. The apparatus according to claim 34, wherein, the second release module includes: a second acquisition sub-module, configured to, in response to detecting a tensor memory release command, acquire header information corresponding to the tensor memory release command; a second release sub-module, configured to release the memory according to the header information.

36. A memory management apparatus, wherein, it includes: a third acquisition module, configured to acquire tensor memory statistical information in training data, the tensor memory statistical information including tensor size, tensor application time, and tensor release time, and the training data including at least one mini-batch data of a deep learning task; a second insertion module, configured to determine a preset traversal order, determine target tensor information according to the preset traversal order, wherein the target tensor information includes target tensor size, target tensor application time, and target tensor release time; determine a target memory address linked list block corresponding to the target tensor size according to the target tensor size, wherein one or more cache blocks corresponding to the target tensor size are provided in the target memory address linked list block, and historical tensor memory statistical information is stored in the cache block; determine whether there is a cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information; when there is the cache block in the target memory address linked list block, insert the target tensor information into the cache block, and determine the next target tensor information for processing according to the preset traversal order; A determination module, configured to determine the current memory address linked list block information as memory allocation information in response to the end of traversal of the tensor memory statistics.

37. The apparatus according to claim 36, wherein, the second insertion module is configured to: when there is no cache block in the target memory address linked list block that does not cause a preset conflict with the target tensor information, traverse the memory address linked list blocks with sizes larger than the target memory address linked list block from small to large, and when there is a cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, insert the target tensor information into the cache block that does not cause a preset conflict with the target tensor information, and determine the next target tensor information for processing according to the preset traversal order. When there is no cache block in the memory address linked list block that does not cause a preset conflict with the target tensor information, create a new cache block in the target memory address linked list block, insert the target tensor information into the new cache block, and determine the next target tensor information for processing according to the preset traversal order.

38. The apparatus according to claim 36, wherein, the part that determines the target memory address linked list block corresponding to the target tensor size according to the target tensor size is configured to: determine the memory address interval of the memory address linked list block; round up the result of dividing the target tensor size by the memory address interval to obtain the required number of memory address intervals; multiply the memory address interval by the required number of memory address intervals to obtain the target memory address size; determine the memory address linked list block corresponding to the target memory address size as the target memory address linked list block.

39. The apparatus according to any one of claims 36-38, wherein, further comprising: a third reallocation module, configured to reallocate the memory reallocation failed tensor according to a preset memory allocation rule in response to detecting a tensor memory allocation failure event.

40. The apparatus according to claim 39, wherein, the third reallocation module includes: a sixth determination sub-module, configured to determine a candidate memory address linked list block corresponding to the memory allocation failed tensor size in response to detecting a tensor memory allocation failure event; a seventh determination sub-module, configured to determine whether there is an idle cache block in the candidate memory address linked list block; a sixth allocation sub-module, configured to allocate the idle cache block to the memory allocation failed tensor and fill the first header information in the idle cache block when there is an idle cache block in the candidate memory address linked list block, wherein the first header information stores the memory address linked list block information where the idle cache block is located; a seventh allocation sub-module, configured to apply for memory allocation from the operating system based on the memory allocation failed tensor size and fill the second header information in the allocated memory when there is no idle cache block in the candidate memory address linked list block, wherein the second header information stores the operating system memory allocation information.

41. An electronic device, It is characterized in that it includes a memory and a processor; wherein the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method steps described in any one of claims 1-20.

42. A computer-readable storage medium, on which computer instructions are stored It is characterized in that when the computer instructions are executed by a processor, the method steps described in any one of claims 1-20 are implemented.

Citation Information

Patent Citations

  • Memory allocation method and device

    CN108874532A

  • Memory management method and device, mobile terminal and storage medium

    CN109815162A