Memory allocation and deallocation decision system supporting dynamic recalculation and method thereof

By dividing the memory pool into sub-memory pools and optimizing memory allocation and release decisions, the problem of excessive memory release in dynamic recomputation is solved, resulting in faster memory acquisition and shorter network training time.

CN115185692BActive Publication Date: 2026-02-17BEIJING SILICON MOBILE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210842115.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2026-02-17
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

Existing technologies suffer from over-release and repeated release of memory space in dynamic recomputation, which leads to extended network training time and fails to effectively meet the memory requirements of large models.

Method used

The memory pool is divided into multiple sub-memory pools by using a memory pool initialization component. The memory allocation and release decisions are optimized by using a memory space monitoring and allocation component and a recomputation cost evaluation component, thereby reducing recomputation overhead and improving the continuity and availability of memory space.

Benefits of technology

By optimizing memory allocation and release decisions, the overhead of heavy computation is reduced, network training time is shortened, and the efficiency and flexibility of memory space utilization are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115185692B_ABST
    Figure CN115185692B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a memory allocation and release decision system supporting dynamic recalculation and a method thereof. The system comprises: a memory pool initialization component applying a continuous memory space as a memory pool and dividing it into at least two sub-memory pools; a memory space monitoring and allocating component polling all sub-memory pools based on a request and allocating the continuous memory space in the first polled sub-memory pool having a continuous memory space greater than the requested size of the current request to the computing logic node making the request; a continuous tensor obtaining component sequentially searching for continuous tensor subsequences having a sum of occupied memory spaces greater than a first size in the memory pool by the Stein-Mehlhorn algorithm when it is determined that there is no continuous memory space greater than the requested size of the current request; a recalculation cost evaluation component calculating the total recalculation cost of the continuous tensor subsequences; and a memory releasing component sorting the total recalculation cost of all continuous tensor subsequences and releasing the memory occupied by all tensors contained in the continuous tensor subsequence corresponding to the smallest total recalculation cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a data processing technology. More specifically, the present disclosure relates to a memory allocation and release decision system supporting dynamic recomputation and a method thereof. BACKGROUND

[0002] In the field of deep learning, the increase of model size and complexity puts higher demands on the size of computing storage resources, especially GPU memory (or memory). In fact, in the first phase of model training, forward calculation, a large number of intermediate variables are saved. A large part of these variables will not be used in a short time after being generated, but will be used in the back propagation phase. Dynamic tensor recomputation (DTR) is based on this phenomenon, which enables the model to automatically release variables in the memory pool during training, and recalculate the values of these variables during back propagation. In this way, the memory occupation is reduced, and the purpose of training large models on limited memory is achieved.

[0003] The core idea of dynamic recomputation is to constantly find the tensor with the minimum calculation cost from the memory pool for release (or false deletion) when the memory occupation exceeds the set threshold, until the memory is below the threshold again. This calculation cost is proportional to the calculation overhead of the tensor. In practice, people tend to recalculate ReLU rather than Conv2d.

[0004] The cost or overhead of recomputation is not only the calculation overhead of the tensor to be released, but also the calculation overhead of the parent and child tensors that are not in the memory. The reason for this is that when a current tensor is recomputed, if its parent tensor is also not in the memory, the parent tensor must be calculated first, and then the current tensor is calculated based on the parent tensor. By including the calculation overhead of the parent and child tensors that are not in the memory in the current tensor's recomputation overhead, this continuous recomputation is inhibited, resulting in the need for repeated memory release processes to meet the demand for larger memory space.

[0005] For example, in the current need to vacate space to accommodate a 100M tensor, generally if the release of a 60M tensor occupied space and then release a 40M tensor occupied space can be placed in the 100M tensor. But if the existing release mode, it is likely that the released 60M tensor and 40M tensor is not continuous placement, that is, 60M tensor and 40M tensor occupied space in the memory pool is not continuous, even if the existing technology to release the 60M tensor and 40M tensor occupied space, still no continuous space to place the 100M memory space, so it is necessary to release again. Obviously this will certainly increase the network training time.

[0006] Therefore, there is a need for a memory allocation and release decision system and method that can support dynamic recomputation, which can avoid excessive repeated release operations to meet the dynamic recomputation memory allocation and release. SUMMARY

[0007] Therefore, in order to obtain the releasable continuous memory space faster on the basis of reducing the recomputation overhead in estimating the recomputation overhead of the tensor, shorten the time of network training, the disclosure provides a memory allocation and release decision system supporting dynamic recomputation, comprising: a memory pool initialization component, which applies a continuous memory space of a predetermined size to a data processing network as a memory pool, and divides the memory pool of the predetermined size into at least two sub-memory pools; a memory space monitoring and allocation component, which is deployed in a memory allocator, determines whether there is a continuous memory space greater than the requested size of the current request in one of all sub-memory pools based on the current request of the current computing logic node in the data processing network requesting allocation of continuous memory space for its data processing, and allocates the continuous memory space in the first polled sub-memory pool with the continuous memory space greater than the requested size of the current request to the computing logic node making the request, starting from the next memory pool of the sub-memory pool allocated by the previous request based on the current request; a continuous tensor acquisition component, which sequentially finds a continuous tensor sub-sequence in the memory pool whose sum of occupied memory space is greater than the first size by the Sieve method in the address order of the memory pool after the memory space monitoring and allocation component polls all sub-memory pools once, and determines that there is no continuous memory space greater than the requested size of the current request in any sub-memory pool; a recomputation cost evaluation component, which calculates the recomputation cost of each current tensor constituting each continuous tensor sub-sequence, and sums up to obtain the total recomputation cost of the continuous tensor sub-sequence; and a memory release component, which sorts the total recomputation cost of all continuous tensor sub-sequences according to the size, and releases the memory occupied by all tensors contained in the continuous tensor sub-sequence corresponding to the smallest total recomputation cost, thereby increasing the allocable memory space of the memory pool.

[0008] According to the memory allocation and release decision system supporting dynamic recomputation of the disclosure, the memory space monitoring and allocation component alternately allocates the requested memory space in the two sub-memory pools in the order of the current request when there are only two sub-memory pools, wherein in the case of sequentially allocating memory from the starting address of the first sub-memory pool, the second sub-memory pool allocates in reverse order from the end address.

[0009] According to the memory allocation and release decision system supporting dynamic recalculation of the present disclosure, when more than two sub-memory pools exist to satisfy the requested memory space size, the memory space monitoring allocation component allocates memory space in the tensor order with the same remainder of the sequence number relative to the modulus n of the number of sub-memory pools, in the sub-memory pool numbered m, where k is an integer coprime with n, where the adjacent two sub-memory pools of the n sub-memory pools are allocated in one from the start address in sequence and the other from the end address in reverse order.

[0010] According to the memory allocation and release decision system supporting dynamic recalculation of the present disclosure, where k is the number closest to n / 2.

[0011] According to the memory allocation and release decision system supporting dynamic recalculation of the present disclosure, when the memory space monitoring allocation component polls the current sub-memory pool in all sub-memory pools and confirms that the remaining memory space is less than the requested size of the current request, it extends the poll to the adjacent sub-memory pool along the polled sub-memory pool, so that when the contiguous memory space at the junction between the current sub-memory pool and its adjacent sub-memory pool is not less than the contiguous memory space of the requested size of the current request, the contiguous space at the adjacent junction obtained by the extended poll is allocated to the computing logic node of the proposed request, and when the contiguous memory space at the junction between the current sub-memory pool and its adjacent sub-memory pool is still less than the contiguous memory space of the requested size of the current request, the polling of the next sub-memory pool is skipped.

[0012] According to the memory allocation and release decision system supporting dynamic recalculation of the present disclosure, the number of tensors contained in the contiguous tensor subsequence is one, two or more than two tensors with contiguous storage space, or two or more than two tensors with free space between their storage spaces.

[0013] According to the memory allocation and release decision system supporting dynamic recalculation of the present disclosure, the recalculation cost evaluation component estimates the forward recalculation cost of each tensor to be estimated as each tensor in the currently saved tensors in the memory pool, where the free space between two tensors is considered as a tensor with zero recalculation cost.

[0014] According to the memory allocation and release decision system supporting dynamic recomputation of the present disclosure, the recomputation cost evaluation component comprises: a tensor usage history recording unit recording tensor usage history and providing a time interval from a latest usage time of a to-be-evaluated tensor to a current time; a tensor recursive query unit recursively querying parent tensors and child tensors of each to-be-evaluated tensor of each continuous tensor sub-sequence in a memory pool, and the recursive query terminates at a first parent tensor and a first child tensor of the to-be-evaluated tensor which are currently existing in the memory pool, thereby determining a first set of all parent tensors and all child tensors for evaluating a recomputation cost of the to-be-evaluated tensor, wherein all parent tensors comprise all intermediate parent tensors between the first parent tensor and the to-be-evaluated tensor, and all child tensors comprise all intermediate child tensors between the first child tensor and the to-be-evaluated tensor; and a cost evaluation unit summing up a forward computation cost of each to-be-evaluated tensor of each continuous tensor sub-sequence and a forward computation cost of each tensor in the first set to obtain a forward computation cost sum of the to-be-evaluated tensor and a total recomputation cost of the continuous tensor sub-sequence.

[0015] According to the memory allocation and release decision system supporting dynamic recomputation of the present disclosure, the recomputation cost evaluation component comprises: a tensor usage history recording unit recording tensor usage history and providing a time interval from a latest usage time of a to-be-evaluated tensor to a current time; a tensor recursive query unit recursively querying parent tensors and child tensors of each to-be-evaluated tensor of each continuous tensor sub-sequence in a memory pool, and the recursive query terminates at a first parent tensor and a first child tensor of the to-be-evaluated tensor which are currently existing in the memory pool, thereby determining a first set of all parent tensors and all child tensors for evaluating a recomputation cost of the to-be-evaluated tensor, wherein all parent tensors comprise all intermediate parent tensors between the first parent tensor and the to-be-evaluated tensor, and all child tensors comprise all intermediate child tensors between the first child tensor and the to-be-evaluated tensor; a cost evaluation unit summing up a forward computation cost of each to-be-evaluated tensor of each continuous tensor sub-sequence and a forward computation cost of each tensor in the first set to obtain a forward computation cost sum of the to-be-evaluated tensor, calculating a recomputation cost time ratio between the forward computation cost sum of the to-be-evaluated tensor and the time interval of the to-be-evaluated tensor, summing up corresponding recomputation cost time ratios of all to-be-evaluated tensors of each continuous tensor sub-sequence to obtain a recomputation cost time ratio sum, and calculating a ratio between the recomputation cost time ratio sum and a spatial size of each continuous tensor sub-sequence as a total recomputation cost of each continuous tensor sub-sequence.

[0016] According to another aspect of the present disclosure, a memory allocation and release decision method supporting dynamic recomputation is provided, comprising: by a memory pool initialization component, applying a continuous memory space of a predetermined size to a data processing network as a memory pool, and dividing the memory pool of the predetermined size into at least two sub-memory pools; based on a current request of a current computation logic node in the data processing network requesting allocation of a continuous memory space for data processing thereof, by a memory space monitoring allocation component deployed in a memory allocator, polling all the sub-memory pools from a next memory pool of a sub-memory pool from which a memory space is allocated based on a previous request of the current request, to determine whether there is a continuous memory space in one of the sub-memory pools that is greater than a requested size of the current request, and allocating the continuous memory space in the first polled sub-memory pool in which there is a continuous memory space greater than the requested size of the current request to the computation logic node that makes the request; when it is determined by the memory space monitoring allocation component that there is no continuous memory space greater than the requested size of the current request in any of the sub-memory pools after polling all the sub-memory pools once, by a continuous tensor obtaining component, sequentially searching for continuous tensor subsequences in the memory pool in which a sum of occupied memory spaces is greater than a first size in an address order of the memory pool by the Sieve method; by a recomputation cost evaluation component, calculating a recomputation cost of each current tensor constituting each continuous tensor subsequence, and summing up to obtain a total recomputation cost of the continuous tensor subsequence; and by a memory release component, sorting total recomputation costs of all continuous tensor subsequences according to size, and releasing all tensors contained in a continuous tensor subsequence corresponding to the smallest total recomputation cost, thereby increasing a redistributable memory space of the memory pool.

[0017] According to the memory allocation and release decision method supporting dynamic recomputation of the present disclosure, when there are only two sub-memory pools, the memory space monitoring allocation component alternately allocates requested memory spaces in the two sub-memory pools in an order of current requests, wherein in a case of sequentially allocating memory spaces from a starting address of a first sub-memory pool, a second sub-memory pool allocates memory spaces in reverse order from a terminal address thereof.

[0018] According to the memory allocation and release decision method supporting dynamic recomputation of the present disclosure, when there are more than two sub-memory pools that satisfy a requested memory space size, based on a sequence number of a current request, in a number n of sub-memory pools, a tensor is sequentially allocated in a sub-memory pool numbered m, where a remainder of the sequence number with respect to the modulus n is 1+(m-1)*k, where k is an integer coprime with n, wherein when allocating memory spaces in adjacent two of the n sub-memory pools, one is sequentially allocated from a starting address, and the other is allocated in reverse order from a terminal address.

[0019] The memory allocation and release decision method supporting dynamic recalculation according to the present disclosure, wherein k is a number closest to n / 2.

[0020] The memory allocation and release decision method supporting dynamic recalculation according to the present disclosure, wherein the memory space monitoring allocation component extends the polling to the adjacent sub-memory pool along the polled sub-memory pool when polling the current sub-memory pool in all sub-memory pools and confirming that the remaining memory space thereof is less than the requested size of the current request, so as to allocate the contiguous space at the adjacent junction obtained by the extended polling to the computing logic node of the proposed request when the contiguous memory space at the adjacent junction between the current sub-memory pool and its adjacent sub-memory pool is not less than the requested size of the current request, and skip the polling of the next sub-memory pool for the current sub-memory pool and its adjacent sub-memory pool when the contiguous memory space at the adjacent junction between the current sub-memory pool and its adjacent sub-memory pool is still less than the requested size of the current request.

[0021] The memory allocation and release decision method supporting dynamic recalculation according to the present disclosure, wherein the number of tensors contained in the contiguous tensor subsequence is one, two or more than two tensors with contiguous storage spaces, or two or more than two tensors with free spaces between the storage spaces.

[0022] The memory allocation and release decision method supporting dynamic recalculation according to the present disclosure, wherein the recalculation cost evaluation component estimates the forward recalculation cost of each tensor to be estimated as each tensor of the tensors currently saved in the memory pool, and regards the tensor with the free space between two tensors as a tensor with zero recalculation cost.

[0023] According to the memory allocation and release decision method supporting dynamic recomputation of the present disclosure, the recomputation cost evaluation performed by the recomputation cost evaluation component includes: recording tensor usage history by a tensor usage history recording unit and providing a time interval from the latest usage time of a to-be-evaluated tensor to the current time; starting from each to-be-evaluated tensor of each continuous tensor sub-sequence in the memory pool, recursively querying the parent and child tensors of the to-be-evaluated tensor by a tensor recursive query unit, and the recursive query terminates at the first parent and child tensors of the to-be-evaluated tensor that are currently present in the memory pool, thereby determining a first set of all parent and child tensors for evaluating the recomputation cost of the to-be-evaluated tensor, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the to-be-evaluated tensor, and all child tensors include all intermediate child tensors between the first child tensor and the to-be-evaluated tensor; and summing the forward computation cost of each to-be-evaluated tensor of each continuous tensor sub-sequence and the forward computation cost of each tensor in the first set by a cost evaluation unit to obtain the forward computation cost sum of each to-be-evaluated tensor and the total recomputation cost of each continuous tensor sub-sequence.

[0024] According to the memory allocation and release decision method supporting dynamic recomputation of the present disclosure, the recomputation cost evaluation performed by the recomputation cost evaluation component includes: recording tensor usage history by a tensor usage history recording unit and providing a time interval from the latest usage time of a to-be-evaluated tensor to the current time; starting from each to-be-evaluated tensor of each continuous tensor sub-sequence in the memory pool, recursively querying the parent and child tensors of the to-be-evaluated tensor by a tensor recursive query unit, and the recursive query terminates at the first parent and child tensors of the to-be-evaluated tensor that are currently present in the memory pool, thereby determining a first set of all parent and child tensors for evaluating the recomputation cost of the to-be-evaluated tensor, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the to-be-evaluated tensor, and all child tensors include all intermediate child tensors between the first child tensor and the to-be-evaluated tensor; and summing the forward computation cost of each to-be-evaluated tensor of each continuous tensor sub-sequence and the forward computation cost of each tensor in the first set by a cost evaluation unit to obtain the forward computation cost sum of each to-be-evaluated tensor and the total recomputation cost of each continuous tensor sub-sequence.

[0025] By adopting the memory allocation and release decision system supporting dynamic recomputation and the method thereof according to the present disclosure, by dividing the memory pool of a predetermined threshold capacity into multiple sub-memory pools, the allocation of continuous tensors in the network computing process is reduced to be continuously distributed in the memory pool, and the tensors in a larger continuous space are released by the Johnson algorithm, which reduces the situation that the continuous tensors are released in the network computing process, resulting in a large recomputation overhead, and also speeds up the time of obtaining the continuous memory space that can be released and shortens the network training time. And by adopting the whole release mode of the continuous space of a predetermined size according to the present disclosure, the memory release decision is more flexible and comprehensive, and the recomputation cost is greatly reduced while achieving the memory release target.

[0026] Other advantages, objects, and features of the application will be understood by those skilled in the art from the following description, and will be appreciated by those skilled in the art from the study and practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 Fig. 1 shows a schematic diagram of a first embodiment of a memory allocation and release decision system supporting dynamic recomputation according to the present disclosure.

[0028] Figure 2 Fig. 2 shows a schematic diagram of an application scenario of memory allocation of a memory allocation and release decision system supporting dynamic recomputation according to the present disclosure.

[0029] Figure 3 Fig. 3 shows a schematic diagram of one example of a scenario of a memory allocation and release decision system 100 supporting dynamic recomputation according to the present disclosure.

[0030] Figure 4 Fig. 4 shows a schematic diagram of another example of a scenario of a memory allocation and release decision system 100 supporting dynamic recomputation according to the present disclosure.

[0031] Figure 5 Fig. 5 shows a schematic diagram of another example of a scenario of a memory allocation and release decision system 100 supporting dynamic recomputation according to the present disclosure.

[0032] Figure 6 Fig. 6 shows a schematic diagram of a flow of a memory release decision method supporting reverse dynamic recomputation according to the present disclosure.

[0033] Figure 7 Fig. 7 shows another application scenario of memory allocation of a memory allocation and release decision system supporting dynamic recomputation according to the present disclosure. DETAILED DESCRIPTION

[0034] The present disclosure will be further described in conjunction with the following examples and drawings, in which like numbers refer to like elements throughout.

[0035] The illustrative examples set forth herein will be described in conjunction with the appended drawings, in which like numbers represent like elements throughout. The following description is not intended to limit the disclosure to one or more particular examples described herein. Rather, the examples set forth herein are intended to illustrate at least some of the various ways in which the disclosure can be practiced.

[0036] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0037] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a particular order or hierarchy among the information. These terms are used merely as labels to identify particular sets of the information. For example, a first tensor could be termed a second tensor without departing from the scope of the present disclosure. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining".

[0038] For a better understanding of the present disclosure, reference will be made to the following descriptions of specific embodiments in connection with the accompanying drawings and specific embodiments.

[0039] Figure 1 A schematic diagram of a first embodiment of a system for memory allocation and deallocation decisions that support dynamic recomputation is shown in accordance with the present disclosure. As Figure 1 As shown, a system for memory allocation and deallocation decisions that support dynamic recomputation 100 in accordance with the present disclosure includes a memory allocator 110, a memory pool initialization component 120, a memory space monitoring and allocation component 130, a contiguous tensor acquisition component 140, a recomputation cost evaluation component 150, and a memory deallocation component 160.

[0040] The memory pool initialization component 120, upon start of processing by the data processing network or upon user initiated recomputation function, applies for a contiguous memory space of a predetermined size as a memory pool for the data processing network and divides the memory pool of the predetermined size into at least two sub-memory pools. As Figure 1As shown, the upper part of the dashed box OP sequence is a data processing network composed of logical computing nodes, and the lower part of the corresponding implementation box OP sequence is the runtime state of the data processing network. The size of the memory space applied for by the application does not exceed the size of the GPU memory running the data processing network.

[0041] Figure 2 As shown is a schematic diagram of an application scenario of memory allocation of a memory allocation and release decision system supporting dynamic recalculation according to the present disclosure. As shown in the figure, the data processing network is composed of logical computing nodes OP, and the runtime state of the data processing network is composed of implementation boxes OP. Figure 2 The memory pool initialization component 120 applies for a memory space of a predetermined size as a memory pool for the data processing network and divides the memory pool of the predetermined size into at least two sub-memory pools when obtaining the initial information of the data processing network or when the user starts the recalculation function. The data processing network is represented by simplified tensors used by each logical computing node OP, such as tensors A, B, C, and D. Although four OP tensors are shown here, there are much more in the data processing network. The memory pool is evenly divided into three groups, i.e., sub-memory pool 1, sub-memory pool 2, and sub-memory pool 3.

[0042] During the application process, the memory space monitoring and allocation component 130 of the memory allocator 110 determines whether there is a continuous memory space greater than the requested size of the current request in one of all the sub-memory pools based on the current request of the current computing logical node in the data processing network to allocate a continuous memory space for its data processing, starting from the next memory pool of the sub-memory pool from which the memory space allocated based on the previous request of the current request, and allocates the continuous memory space in the sub-memory pool in which the continuous memory space greater than the requested size of the current request is first polled to the computing logical node making the request. Specifically, as shown in the figure, Figure 2As shown, the memory space monitoring and allocating component 130 will first allocate memory space in sub-memory pool 1 for tensor A (if there is enough free memory space in sub-memory pool 1 for tensor A), then allocate memory space in sub-memory pool 2 for tensor B (if there is enough free memory space in sub-memory pool 2 for tensor B), and so on, for tensor C (if there is enough free memory space in sub-memory pool 3 for tensor C), and for tensor D (if there is enough free memory space in sub-memory pool 1 for tensor D) after tensor A, and so on. Alternatively, if there is not enough free memory space in sub-memory pool 2 for tensor B, then the memory space monitoring and allocating component 130 will allocate free memory space in sub-memory pool 3 for tensor B, and if there is still not enough free memory space, then the memory space monitoring and allocating component 130 will allocate free memory space in sub-memory pool 1 for tensor B. If the memory space monitoring and allocating component 130 has traversed through each of the sub-memory pools and there is still no continuous memory space available for the logical computing node that is currently requesting memory space for the current tensor, then the memory space monitoring and allocating component 130 will trigger the query determination process for the release object in the memory releasing process.

[0043] Therefore, when the continuous tensor obtaining component 140 determines that there is no sub-memory pool with a continuous memory space larger than the requested size of the current request after the memory space monitoring and allocating component 130 polls all the sub-memory pools, the continuous tensor obtaining component 140 sequentially searches for a continuous tensor sub-sequence in the memory pool with a sum of occupied memory spaces larger than the first size K by the way of the Sieve method according to the address order of the memory pool. Specifically, the continuous tensor obtaining component 140 arranges all the tensors in the memory pool into a list according to their address order in the memory pool, and sorts them in the list, in which each independent free space is regarded as a free tensor, the calculation cost of the free tensor is zero, and the size of the occupied space of the free tensor is the size of the free space. Then, the continuous tensor obtaining component 140 sequentially selects a continuous tensor sub-sequence with a size larger than the first size K according to the Sieve method. The Sieve method of the present disclosure initializes the start and end points of two indexes of a sieve to 0, and then constantly increases the value of the end point until the size of the continuous tensors covered by the sieve is larger than the first size K. At this time, if the start point is added by 1, and the size of the continuous tensors covered between the start point and the end point after the addition is smaller than the first size K, the value of the start point is kept unchanged, and the value of the end point is increased again until the size of the continuous tensors covered between the start point and the end point of the sieve is larger than the first size K. It should be noted that the addition of 1 here is to move the address of the start point or the end point to the end address of a tensor, that is, to move the address sequence of the memory space occupied by a tensor according to the arrangement order of the tensors. The interval of the address sequence moved by the addition of 1 each time is not the same, but consistent with the size of the tensor. The foregoing process is repeatedly repeated until the entire memory pool is traversed. A series of continuous tensor sub-sequences can be obtained by the Sieve method. The Sieve method is performed in a conventional manner in the art. Since the present disclosure does not make any improvement to the Sieve method, it will not be described in detail here.

[0044] Then, after obtaining the continuous tensor sub-sequence, the recalculation cost evaluating component 150 calculates the recalculation cost of each current tensor constituting each continuous tensor sub-sequence, and sums them up to obtain the total recalculation cost of the continuous tensor sub-sequence. It should be noted that if there is a free space between two tensors in each tensor sub-sequence, the free space is regarded as a tensor with a recalculation cost of zero. The recalculation cost of each current tensor in each tensor sub-sequence can be directly obtained from the calculation cost of the logical calculation unit thereof. Therefore, in the simple mode, the total recalculation cost of the continuous tensor sub-sequence can be obtained by simply summing up the recalculation cost of each current tensor in each tensor sub-sequence.

[0045] Alternatively, since each tensor is not necessarily directly available from its parent tensor or child tensor, because its parent tensor or child tensor can have been released from the memory pool, it is necessary to recursively query for the parent tensor or child tensor that can perform the re-computation to obtain the current tensor, and therefore the re-computation cost of the parent tensor and child tensor closest to the current tensor not existing in the memory pool also needs to be considered. Therefore, the re-computation cost evaluation component 150 estimates the forward re-computation cost of each tensor as the tensor to be estimated, for each of the tensors currently saved in the memory pool. The re-computable forward re-computation cost of each current tensor t in each tensor sub-sequence is calculated by the re-computation cost evaluation component 150 using the following formula:

[0046] h DTR (t) = c(t) / [m(t) s(t)]

[0047] where h DTR (t) is the forward re-computation cost of the current tensor t, m(t) is the space size of the memory pool occupied by the current tensor t, s(t) is the time distance between the time of the last use (i.e., the most recent use) of the current tensor t and the current time, and c(t) is

[0048]

[0049] c(t) represents the sum of the forward computation costs of the current tensor t (i.e., the selected tensor that can be re-computed) and the parent and child tensors t' of the current tensor t not existing in the video memory or the memory pool when re-computation is performed. Wherein, e * (t) represents the set of parent and child tensors t' of the current tensor t not existing in the video memory or the memory pool, and c0(t') is the computation cost of each tensor t' in the set e * (t). e * (t) represents that the elements in the set of parent and child tensors t' of the current tensor t not existing in the video memory or the memory pool are obtained by the re-computation cost evaluation component 150 using recursive querying, i.e., querying the continuous parent and child tensors t' of the current tensor t not existing in the video memory or the memory pool, and if the closest parent or child tensor is found, the recursive querying is stopped. Thus, all parent tensors t' not existing in the memory pool between the parent tensor t' found in the memory pool and the current tensor t, and all child tensors t' not existing in the memory pool between the child tensor t' found in the memory pool and the current tensor t are the set e *The elements in (t) are all the tensors that need to be recomputed when the current tensor t is recomputed in the forward direction. In addition, under the assumption that the current tensor t is released, its descendant tensors that are currently in the memory pool can also be recomputed, and these computation costs are considered as the “sunk costs” of the current tensor t recomputation and should be included in the computation of the current tensor t recomputation cost. Therefore, although h DTR (t) is the forward recomputation cost of the current tensor t, the computation of c(t) also includes the computation cost of the descendant tensors t' of the current tensor t. Therefore, the recomputation cost evaluation component 150 also needs to use recursive queries to obtain all the descendant tensors t' of the current tensor t that do not exist in the memory pool and use them to estimate the recomputation cost of the current tensor t.

[0050] The forward recomputation cost h DTR (t) of each current tensor in the continuous tensor subsequence is summed to obtain the total recomputation cost of each continuous tensor subsequence.

[0051] Alternatively, in the case where the to-be-estimated tensor or the current tensor in the continuous tensor subsequence can be obtained by reverse recomputation, in the present disclosure, in addition to considering the forward recomputation cost of the current tensor t, the reverse recomputation cost of the current tensor t also needs to be considered. The reverse recomputation cost is also estimated, and the minimum value between the forward recomputation cost and the reverse recomputation cost is taken as the current recomputation cost of the to-be-estimated tensor. For the same current tensor t, there will be a difference between the reverse recomputation cost and the forward recomputation cost. Therefore, taking the smaller cost between the two will result in a smaller recomputation cost of the current tensor t. Therefore, the recomputation cost evaluation component 150 can directly use the set of parent and descendant tensors t' of the current tensor t that are not in the video memory or the memory pool obtained by recursive queries to calculate the sum of the reverse computation costs as the reverse impact computation cost, that is,

[0052] h' DTR (t) = c'(t) / [m(t) s(t)]

[0053] where h' DTR (t) is the reverse recomputation cost of the current tensor t, m(t) is the space size of the memory pool occupied by the current tensor t, s(t) is the time distance between the time when the current tensor t was last used (i.e., the most recently used) and the current time, and c'(t) is

[0054]

[0055] c'(t) represents the sum of the reverse computation cost of the current tensor t (i.e. the selected tensor to be re-computed) and the parent and child tensors t' of the current tensor t not in the video memory or the memory pool when re-computation is performed. Wherein, e'(t) represents the set of the parent and child tensors t' of the current tensor t not in the video memory or the memory pool, and c'(t') is the reverse computation cost of each tensor t' in the set e'(t). ( t) represents the set of the parent and child tensors t' of the current tensor t not in the video memory or the memory pool, and c'(t') is the reverse computation cost of each tensor t' in the set e'(t). * t) represents the set of the parent and child tensors t' of the current tensor t not in the video memory or the memory pool, and c'(t') is the reverse computation cost of each tensor t' in the set e'(t).

[0056] The re-computation cost evaluation component 150 estimates the reverse re-computation cost h'(t) of the current tensor t after the forward re-computation cost h0(t) of the current tensor t is estimated, and takes the minimum value of the forward re-computation cost h0(t) and the reverse re-computation cost h'(t) as the current re-computation cost of the tensor to be estimated. The re-computation cost evaluation component 150 sums the current re-computation cost of each current tensor in a continuous tensor sub-sequence to obtain the total re-computation cost of the continuous tensor sub-sequence. DTR t) and the reverse re-computation cost h'(t) as the current re-computation cost of the tensor to be estimated. The re-computation cost evaluation component 150 sums the current re-computation cost of each current tensor in a continuous tensor sub-sequence to obtain the total re-computation cost of the continuous tensor sub-sequence. DTR t) and the reverse re-computation cost h'(t) as the current re-computation cost of the tensor to be estimated. The re-computation cost evaluation component 150 sums the current re-computation cost of each current tensor in a continuous tensor sub-sequence to obtain the total re-computation cost of the continuous tensor sub-sequence. DTR t) and the reverse re-computation cost h'(t) as the current re-computation cost of the tensor to be estimated. The re-computation cost evaluation component 150 sums the current re-computation cost of each current tensor in a continuous tensor sub-sequence to obtain the total re-computation cost of the continuous tensor sub-sequence.

[0057] However, in the case of reverse re-computation, since there are cases that the current tensor t can be re-computed reversely or the logical computation node generating the current tensor t is not reversible (i.e. the current tensor t cannot be re-computed reversely to obtain the parent tensor associated therewith as a child tensor), and the current tensor t cannot be re-computed reversely, the cost of the forward re-computation of the current tensor t needs to be reconsidered. For this purpose, the re-computation cost evaluation component 150 maintains the calculation manner of the cost of the forward re-computation of the current tensor t in the case that the current tensor t can be re-computed reversely or the logical computation node generating the current tensor t is not reversible.

[0058] However, in the case that the current tensor t cannot be re-computed reversely, it is necessary to consider whether there is a continuously re-computable parent tensor in the memory pool to enable the forward re-computation of the current tensor t if the current tensor t is released, so as to enable the forward re-computation of the current tensor t. Therefore, the cost of the reverse re-computation of the parent tensor of the current tensor t needs to be considered at this time. At this time, the actual re-computation cost of the forward re-computation of the current tensor t will be proportionally increased. For this purpose, the re-computation cost evaluation component 150 assigns a cost coefficient coeff(t) to the already calculated h'(t). DTR t) on the basis of the cost coefficient coeff(t),

[0059] coeff(t) = (∑ t′∈r(t) p(t')) / c0(t)

[0060] where r(t) represents the set of parent tensors t' that are recursively queried from the current tensor t by the recompute cost evaluation component 150 in the case that the current tensor t cannot be inversely recomputed, and p(t') is the recompute cost of each parent tensor t' in the set, as follows:

[0061] p(t') = c0(t') + c user (t')

[0062] where c0(t') is the forward computation cost of the parent tensor t', and c user (t') is the sum of the forward recomputation cost of the child tensors t" that are recursively queried from the parent tensor t' and cannot be inversely recomputed. The adjustment coefficient, coeff(t), is obtained by summing p(t') of all parent tensors t' in the set r(t) and calculating the ratio of the sum value to c0(t), and the forward recomputation cost c fwd (t) of the current tensor t in the case that it cannot be inversely recomputed is obtained as follows:

[0063] c fwd (t) = coeff(t) * h DTR (t)

[0064] Alternatively, in the present disclosure, the inverse recomputation cost of the same current tensor t is considered in the case that the inverse computation naturally exists in the backpropagation, and thus the cost of the inverse recomputation of some current tensors t will naturally decrease. To this end, when the recompute cost evaluation component 150 recursively queries the parent and child tensors of the current tensor t, the recursive query process only needs to end when a tensor that is dependent on the backpropagation (such as the input tensor of convolution, BN, and fully connected layer, etc.) is encountered. The reason for this further limitation of the end condition of the recursive query is that, for example, if there is a logical computation node (OP f) in the data processing network, and the input and output tensors A and B of the logical computation node are both dependent on the backpropagation, since the direction of the backpropagation is opposite to the direction of the network itself, B is often used first and then A in the backpropagation. Therefore, it is reasonable to assume that, when A needs to be used, B has not been kicked out of the memory pool because it has just been used in the usual data processing process. Therefore, this means that even if the tensor B does not exist in the memory pool, it is considered to actually exist in the memory pool because it is dependent on the backpropagation, and thus it is considered to be a tensor that is encountered in the recursive query.

[0065] Therefore, although the recomputation cost evaluation component 150 appears to use the set of parent and child tensors t′ obtained by the recursive query that the current tensor t is not in the video memory or memory pool to calculate the sum of the reverse computation costs as the reverse impact computation cost, during the calculation of the reverse recomputation cost, the set e′ of the parent and child tensors t′ that the current tensor t is not in the video memory or memory pool... * (t) and set e * (t) are not the same:

[0066] h′ DTR (t)=c′(t) / [m(t)·s(t)]

[0067] Where h′ DTR (t) represents the cost of recompiling the reverse tensor t, m(t) represents the size of the memory pool occupied by the current tensor t, s(t) represents the time distance between the last time the current tensor t was used (i.e., the most recent time it was used) and the current time, and c′(t) represents the cost of recompiling the reverse tensor t.

[0068]

[0069] c′(t) represents the sum of the reverse computation costs of the current tensor t (i.e., the tensor selected for recomputation) and its parent and child tensors t′ that are not in video memory or the memory pool when recomputing. Here, c′0(t′) is the set e′. * The cost of reverse computation for each tensor t′ in (t), e′ * (t) represents the set of parent and child tensors t′ of the current tensor t that are not in the video memory or memory pool, where the parent and child tensors t′ are the parent and child tensors t′ of the current tensor t and the tensors that are encountered by the recursive query and depended upon by backpropagation that are not in the video memory or memory pool. The set e′ * (t) and set e * (t) Both are the same when the tensor that backpropagation depends on is exactly a tensor that exists in the memory pool, or when the tensor that the recursive query first encounters in the memory pool is not a tensor that backpropagation depends on. However, they are different when the tensor that the backpropagation depends on first is not a tensor that the memory pool first encounters.

[0070] The recomputation cost evaluation component 150 estimates the forward recomputation cost of each tensor currently stored in the memory pool as a tensor to be estimated, and, if the tensor to be estimated can be obtained through reverse recomputation, also estimates its reverse recomputation cost. The minimum of the forward and reverse recomputation costs is taken as the current recomputation cost of the tensor to be estimated. Therefore, the recomputation cost evaluation component 150 for h′... DTR (t) and hDTR (t) performing a min operation:

[0071] min{c fwd (t),c bwd (t)}

[0072] and the result is the re-computation cost of the current tensor t. It is noted that the estimation is performed at any time. In the next round of estimation, if the current tensor t is still in the memory pool, it will be re-estimated, and the results of the two estimations can be different due to the different existing tensors in the memory pool.

[0073] The process of finding the optimal tensor in each round of the forward calculation of the data processing network can be completed by executing the above steps through the following code:

[0074]

[0075] Alternatively, the re-computation cost evaluation component 150 can also obtain the re-computation cost of each continuous tensor sub-sequence by estimating the entire continuous tensor sub-sequence. Specifically, one method is to directly calculate the c / s value of each current tensor in the continuous tensor sub-sequence, where c is the current re-computation cost of each current tensor, and s is the time interval between the current time of each current tensor and the nearest time when it was last used, as described above. After obtaining the c / s value of each current tensor, sum it to obtain the sum of the ratios. At the same time, sum the space occupied by the continuous tensor sub-sequence to obtain the sum of the continuous tensor sub-sequence ratios sum(c / s), i.e., obtain the space size occupied by the continuous tensor sub-sequence, i.e., sum(m). Finally, by calculating the ratio between the sum of the continuous tensor sub-sequence ratios and the space size sum(m), the total re-computation cost of each continuous tensor sub-sequence is obtained.

[0076] Alternatively, another method is that the re-computation cost c of each current tensor in each continuous tensor sub-sequence can obtain the re-computation costs of its parent and child tensors as part of the calculation cost of the current tensor by the recursive query method as described above. For the recursive calculation of the cost of each current tensor, please refer to the above description, which will not be repeated here. After calculating the c of each current tensor by the recursive query method, the total re-computation cost of the continuous tensor sub-sequence is obtained by the sum(c / s) / sum(m) total re-computation cost method as described above.

[0077] After obtaining the total re-computation cost of the continuous tensor sub-sequence, the memory releasing component 160 sorts the total re-computation cost of all continuous tensor sub-sequences according to the size, and releases the memory occupied by all tensors contained in the continuous tensor sub-sequence corresponding to the smallest total re-computation cost, thereby increasing the allocable memory space of the memory pool.

[0078] Referring back to Figure 1 The re-computation cost evaluation component 150 can include a tensor usage history recording unit 151, a tensor recursive query unit 152, and a cost estimation unit 153. The tensor usage history recording unit 151 records the tensor usage history and provides the time interval between the current time and the latest usage time of the current tensor in each continuous tensor sub-sequence as the tensor to be estimated. The tensor recursive query unit 152 recursively queries the parent and child tensors of each tensor to be estimated from the memory pool, and the recursive query terminates at the first parent tensor and the first child tensor of the tensor to be estimated that are currently present in the memory pool, thereby determining the first set of all parent tensors and all child tensors for estimating the re-computation cost of the tensor to be estimated, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the tensor to be estimated, and all child tensors include all intermediate child tensors between the first child tensor and the tensor to be estimated.

[0079] The cost estimation unit 150 is then used to sum the forward computation cost of the tensor to be estimated and the forward computation cost of each tensor in the first set to obtain the forward computation cost sum of the tensor to be estimated, sum the reverse computation cost of the tensor to be estimated and the reverse computation cost of each tensor in the first set to obtain the reverse computation cost sum of the tensor to be estimated, and multiply the size of the space occupied by the tensor to be estimated and the time interval, thereby calculating the ratio of the forward computation cost sum and the product as the forward re-computation cost of the tensor to be estimated, and calculating the ratio of the reverse computation cost sum and the product as the reverse re-computation cost of the tensor to be estimated.

[0080] In the scenario of using the memory allocation and release decision system 100 supporting dynamic re-computation of the present disclosure, for example, a straight-line data processing op sequence: x1->x2->x3->x4->x5, if x1 is a tensor input from the outside world, therefore its logical computation node (op) is empty and cannot be released. The logical computation nodes of x2->x3, x3->x4, and x4->x5 are reversible. Suppose the logical computation nodes of x1->x2, x2->x3, and x3->x4 are relatively complex (such as convolution), and the logical computation node of x4->x5 is relatively simple (such as scalar).

[0081] If x6 is generated from x5 according to the irreversible logic calculation node, and it is assumed that no space can be applied for x6 due to the memory pool threshold limit, then the memory release decision system 100 releases the memory space occupied by the optimal tensor x4 to make space for tensor x6. Because under the above assumption, the tensor x4 is released, and the recalculation cost of the tensor x4 is

[0082] cost(x4) = min(c fwd (x4), c bwd (x4)) = c bwd (x4) = x c5 -- > x4 / m(x4) / s(m4)

[0083] The inverse calculation cost of x5->x4 (x x5->x4 ) is smaller than the calculation cost of recalculating the tensors x2 or x3 after releasing the tensors x2 or x3, and is equivalent to the calculation cost of recalculating the tensor x5 after releasing the tensor x5. In the case of equivalence, if the interval times of the two are different, the interval times of the two are compared. In the case of s(x4) > s(x5), cost(x4) < cost(x5), and therefore the tensor x4 is released.

[0084] In the above case, when x7 is generated, the tensor needs to be released again, and the calculation consumption of the tensors x2 and x3 is large and is not considered. Therefore, the tensor is selected from x5 and x6. If Because s(x5) > s(x6), the tensor x5 is selected to be released without considering the coefficient coeff.

[0085] If the calculation logic nodes of x2->x5 are all reversible, x2->x4 can be inversely recalculated from x5 by O(1) space complexity during back propagation, so it is desired to retain x5 and release x6. Therefore, it is necessary to calculate the coefficients coeff(x5) and coeff(x6):

[0086] coeff(x5) = (p(x2) + p(x3) + p(x4)) / c0(x5)

[0087] = (c0(x2) + c0(x3) + c0(x4)) / c0(x5) > > 1

[0088] And

[0089] coeff(x6) = 1

[0090] Therefore, the cost estimation unit 153 can estimate:

[0091] cost(x5) > cost(x6)

[0092] Thus, the memory release component 160 ensures that the DTR policy releases x6 instead of x5 as intended.

[0093] Figure 3 A schematic diagram of one example of a scenario employing the disclosed memory allocation and release decision system 100 that supports dynamic recomputation is shown. As Figure 3 As shown, the straight-line op sequence: x1 ->... -> x11, which if x1 is a tensor inputted by the outside world, thus its logical computation nodes (ops) are empty and cannot be released. It is assumed that only tensors x5, x8 are left in the current GPU memory pool, and one of the tensors x5 and x8 needs to be released to free up GPU memory. The dashed lines in the figure represent that the logical computation nodes (OPs) of x2 -> x3, x3 -> x4, x4 -> x5 are reversible. Thus, since there are continuous reversible logical computation nodes of the parent tensors of the current tensor t x5, the coefficients coeff(x5) and coeff(x8) of forward recomputation need to be considered:

[0094]

[0095]

[0096] For ease of understanding, the conditions of this example are simplified. It is assumed that c0(x i ) = c, m(x i ) = m, 2≤i≤11.

[0097] Since the time interval s(x5) > s(x8) is used under the straight-line sequence example, without considering the coefficients (not belonging to the case represented by the dashed line in Figure 3 cost(x5) < cost(x8), thus tensor x5 will be released.

[0098] However, in the case where the logical computation nodes (OPs) of x2 -> x3, x3 -> x4, x4 -> x5 are reversible, i.e., x2, x3, x3, x4 can be obtained by inverse computation through x5, thus, if x5 is released in this case without considering the addition of the coefficients employed by the disclosure, this result will cause the excellent property that x2 -> x4 can be recomputed with O(1) space complexity when backpropagating to be discarded, and this loss will also increase with the growth of “x2 ->... -> x4” (i.e., the number of parent nodes of this continuous reversible computation increases). Thus, in addition to considering the estimated cost of inverse recomputation of the tensor, the disclosure also considers the case where the excellent property loss is eliminated by adding the coefficients of forward recomputation.

[0099] To this end, in Figure 3In the illustrated example, as assumed above, in the case of c0(x i ) = c,m(x i ) = m,2≤i≤11, coeff(x5) = 3; coeff(x8) = 1, so cost(x5) > cost(x8), thus the memory release component 160 will preferentially release the tensor x8.

[0100] Figure 4 Another example of a scenario employing the memory allocation and release decision system 100 of the present disclosure to support dynamic recomputation is illustrated in the schematic diagram shown in FIG. 6. As shown in FIG. 6, the data processing network is in a sequence scenario where it is assumed that only tensors x5, y1 remain in the GPU pool. The op x2->x3->x4->x5 is reversible. Figure 4

[0101] According to the above-described recomputation cost estimation approach of the present disclosure, the forward recomputation cost of tensors x5, y1 is calculated as follows:

[0102]

[0103]

[0104] coeff(x5) = (p(x2) + p(x3) + p(x4)) / c0(x5)

[0105] = (c0(x2) + c0(x3) + c0(x4) + c0(y1) + c0(y2)) / c0(x5)

[0106] ≈5

[0107] coeff(y1) = 1

[0108] If the computation consumption and GPU memory occupation of this computation sequence are comparable, then the cost(y1) and cost(x5) are close in size without considering the coefficients. However, because there are more continuous reversible logical computation nodes before x5, it is more valuable to retain the tensor x5 than the tensor y1. Because the tensor y1 is not reversible, it can be obtained by forward recomputation from x4, thus after releasing the tensor y1, the tensor y1 can be recomputed by x5->x4->y1, but after releasing the tensor x5, x5 can only be recomputed by x1->x2->...->x5, which consumes more and has a greater recomputation cost. Thus, by recursively querying the continuity of the reversible parent tensors of the current tensor t, a higher recomputation cost coefficient is assigned, so that the current tensor with a continuous reversible parent tensor is retained to prevent it from being released and requiring an excessively high recomputation cost.

[0109] Figure 5 ​Illustrated is a schematic diagram of another example of a scenario employing the disclosed memory allocation and deallocation decision system 100 that supports dynamic recomputation. Figure 5 The illustrated example takes a reversible residual network block (RevNet Block) as the basic unit, where X1->Y1->Y2->Y3 are 3 consecutive and reversible residual blocks.

[0110] According to the design of RevNet, the calculation of tensors X1->Y1->Y2->Y3 is reversible, so in the forward calculation process, X1, x11, x12, y11, y12, y21, y22, y31, y32 can be released, only Y3 is retained, and in the back propagation, Y3 and its gradient (grad) are used to calculate the previous items and their gradients in turn. In line with the design of RevNet, the DTR strategy of adding reverse recomputation cost cost and coefficient also gives priority to retaining tensors such as Y3 at the end of the continuous reversible calculation chain. Since the reversible modules are connected in series, the dimensions of X1 and Y3 are the same, the dimensions of x11, x12, y11, y12, y21, y22, y31, y32 after blocking are the same, and the number of channels is half of X1 and Y3. Therefore, the actual occupied memory space of each is as follows:

[0111]

[0112] The time interval relationship between the most recent access or use time S of these tensors is as follows:

[0113] s(X1)>s(x 11 )>s(x 12 )>…>s(y 31 )=s(y 32 )>s(Y3)

[0114] Therefore, at the time of initial release, the neighbor cost of all tensors is zero (neighbor_cost=0). Assuming that the logical node function (function F, G) structure is similar, the internal OP reversibility is unknown (treated as irreversible), then there is a relationship between the calculation cost or calculation consumption as follows:

[0115] c chunk ≈c concat ≈c add <<c op in F,G

[0116] In this case, it is generally desirable that the computation cost of tensor Y3 is less than the computation cost of the tensors represented by lower case letters and the intermediate variables in the F, G modules. The reason is that as long as Y3 is kept, any one of the above mentioned variables can be computed reversely. However, the tensors with small computation cost are usually released first. Therefore, if Y3 is released, it will be difficult to compute reversely the tensors in front of it. Therefore, considering that there are parent tensors that can be computed reversely in succession, it is desirable to keep Y3 as long as possible instead of releasing it. This is the reason why the system is applied in the present disclosure, that is, the computation cost of the parent tensors that can be computed reversely in succession is counted in the computation cost of the current tensor t.

[0117] If no coefficient is added, cost(Y3) = c concat / m(Y3) / s(Y3) will be very small (because the computation of concat consumes very little). In this case, Figure 5 The entire sequence of tensor computation costs shown above makes the order of the released tensors between the tensors roughly X1, Y3, y11-y22, and finally the intermediate variables in F, G. According to formula (1)

[0118]

[0119] r(Y3) = {X1, x 12 , x 11 , y 12 , y 11 , y 22 , y 21 , y 32 , y 31}

[0120] p(X1) = c0(X1)

[0121] p(x 11 ) = c chunk

[0122]

[0123]

[0124]

[0125] p(y 32 ) = c add

[0126] cost(Y3) = coeff(Y3) · c concat / m(Y3) / s(Y3)

[0127]

[0128] where t e r(Y3) U F U G

[0129] As can be seen from the calculation of cost(Y3), the numerator part contains the calculation cost of the set r(Y3) and the calculation cost of all tensors in F and G, so the re-computation cost of the tensor Y3 will be the largest, which leads to that the tensor Y3 will be released the latest in the tensor sequence, or it is the least likely to be selected as the released tensor.

[0130] Alternatively, as shown in Figure 1 If the one parent tensor or one child tensor of the to-be-evaluated tensor to be used by the back propagation is obtained by the recursive query before the first parent tensor and the first child tensor of the to-be-evaluated tensor currently existing in the memory pool are obtained by the recursive query, the one parent tensor or one child tensor to be used by the back propagation is determined as the first parent tensor or the first child tensor.

[0131] Alternatively, the tensor recursive query unit 152 obtains, by recursive query, a second tensor set of the successively irreversibly computable parent tensors of the to-be-evaluated tensor and the respective successively non-reversibly computable child tensors of the successively irreversibly computable parent tensors in the case that the to-be-evaluated tensor cannot be irreversibly re-computed based on the tensors stored in the current memory pool or the calculation logic node itself of the to-be-evaluated tensor is itself irreversibly, based on the information recorded by the tensor usage history recording unit 151; and the cost estimation unit 153 sums the forward computation costs of the tensors in the second tensor set to obtain a cost correction coefficient, and uses the product of the cost correction coefficient and the forward re-computation cost as the forward re-computation cost of the to-be-evaluated tensor.

[0132] Alternatively, as shown in Figure 1 The memory space monitoring and allocating component 130 also monitors whether any continuous allocable memory space in the memory pool is less than the memory space applied by any calculation logic node to the memory pool. The re-computation cost evaluation component 150 also estimates, for each of the tensors currently saved in the memory pool, the forward re-computation cost of each tensor as a to-be-evaluated tensor, and in the case that the to-be-evaluated tensor can be obtained by reverse re-computation, also estimates the reverse re-computation cost thereof, and takes the minimum value of the forward re-computation cost and the reverse re-computation cost as the current re-computation cost of the to-be-evaluated tensor when any continuous allocable memory space in the memory pool is less than the memory space applied by any calculation logic node to the memory pool.

[0133] Figure 6 As shown in the flowchart of the memory release decision method supporting reverse dynamic re-computation according to the present disclosure. AsFigure 6 As shown, first, at step S610, the memory pool initialization component 120 applies a continuous memory space of a predetermined size as a memory pool for the data processing network and divides the memory pool of the predetermined size into at least two sub-memory pools. Subsequently, when the data processing network is in a runtime state, at step S640, based on a current request of a current computing logic node in the data processing network requesting allocation of a continuous memory space for its data processing, the memory space monitoring allocation component 130 deployed in the memory allocator 110 polls all the sub-memory pools starting from the next sub-memory pool of the sub-memory pool from which a memory space is allocated based on a previous request of the current request, to determine whether there is a continuous memory space in one of the sub-memory pools that is greater than the requested size of the current request, and allocates the continuous memory space in the first polled sub-memory pool in which there is a continuous memory space greater than the requested size of the current request to the requesting computing logic node. If, at step S640, after the memory space monitoring allocation component 130 polls all the sub-memory pools once, it is determined that there is no continuous memory space in any of the sub-memory pools that is greater than the requested size of the current request, then at step S630, the continuous tensor acquisition component 140 sequentially finds continuous tensor subsequences in the memory pool in which the sum of occupied memory spaces is greater than the first size by the Sieve of Eratosthenes according to the address order of the memory pool. Next, at step S640, the recalculation cost evaluation component 150 calculates the recalculation cost of each current tensor constituting each continuous tensor subsequence and sums them up to obtain the total recalculation cost of the continuous tensor subsequence. Finally, at step S650, the memory release component 160 sorts the total recalculation cost of all continuous tensor subsequences according to size, and releases the memory occupied by all tensors contained in the continuous tensor subsequence corresponding to the smallest total recalculation cost, thereby increasing the redistributable memory space of the memory pool.

[0134] Alternatively, the recalculation cost evaluation step S640 can be performed by summing up the computation costs of each tensor in the continuous tensor subsequence as described above, or by summing up the forward or backward recalculation costs of each tensor in the continuous tensor subsequence, or by summing up the recalculation costs of each tensor of the recursive parent and child tensors of each tensor and then performing the summation. In the case of considering the recalculation costs of the recursive parent and child tensors as described above, first, in any data processing process, the input, output and generated intermediate tensors of each data computation logic node entering the runtime are recorded. This way of recording the tensor usage history is the existing way, and thus is not described in detail in the disclosure. Therefore, at step S641, the time interval S from the last usage time of the to-be-evaluated tensor to the current time is obtained by the tensor usage history recording unit 151. Then, at step S642, the parent and child tensors of each to-be-evaluated tensor are recursively queried from the memory pool by the tensor recursive query unit 152, and the recursive query terminates at the first parent tensor and the first child tensor of the to-be-evaluated tensor which are currently present in the memory pool, thereby determining the first set of all parent tensors and all child tensors for evaluating the recalculation cost of the to-be-evaluated tensor, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the to-be-evaluated tensor, and all child tensors include all intermediate child tensors between the first child tensor and the to-be-evaluated tensor. That is, the set of continuous parent and child tensors which are not present in the memory pool from the to-be-evaluated tensor forward and backward. Then, based on these tensor sets, at step S643, the forward computation cost of the to-be-evaluated tensor and the forward computation costs of each tensor in the first set are summed up by the cost evaluation unit 153 to obtain the forward computation cost sum of the to-be-evaluated tensor, the backward computation cost of the to-be-evaluated tensor and the backward computation costs of each tensor in the first set are summed up to obtain the backward computation cost sum of the to-be-evaluated tensor, and the size of the space occupied by the to-be-evaluated tensor and the time interval are multiplied to obtain the product, thereby calculating the ratio of the forward computation cost sum and the product as the forward recalculation cost of the to-be-evaluated tensor, and calculating the ratio of the backward computation cost sum and the product as the backward recalculation cost of the to-be-evaluated tensor.

[0135] Alternatively, at step S642, if one parent tensor or one child tensor of the to-be-estimated tensor to be used by back propagation is found by the tensor recursive query unit 152 before the first parent tensor and the first child tensor of the to-be-estimated tensor currently existing in the memory pool are recursively queried, the one parent tensor or the one child tensor to be used by back propagation is determined as the first parent tensor or the first child tensor, and the recursive query is terminated.

[0136] Alternatively, at step S642, if it is learned based on the information recorded by the tensor usage history recording unit that the to-be-estimated tensor cannot be reversibly recomputed or generated based on the tensors stored in the current memory pool, or the calculation logic node of the to-be-estimated tensor itself is reversible, the tensor recursive query unit 152 recursively queries a second tensor set of the parent tensors of the continuous reversible reverse computation of the to-be-estimated tensor and the respective child tensors of the continuous irreversible reverse computation of the parent tensors of the continuous reversible reverse computation. Then at step S643, the cost estimation unit sums the forward computation costs of each tensor in the second tensor set to obtain a cost correction coefficient, and uses the product of the cost correction coefficient and the forward recomputation cost as the forward recomputation cost of the to-be-estimated tensor.

[0137] Alternatively, at step S610, the memory space monitoring and allocation component further monitors whether any continuous allocable memory space of the memory pool is less than the memory space applied by any calculation logic node to the memory pool. And at step S640, when any continuous allocable memory space of the memory pool is less than the memory space applied by any calculation logic node to the memory pool, the recomputation cost evaluation component estimates the forward recomputation cost of each tensor as the to-be-estimated tensor for each tensor currently saved in the memory pool, and estimates the reverse recomputation cost thereof if the to-be-estimated tensor can be obtained by reverse recomputation, and takes the minimum value of the forward recomputation cost and the reverse recomputation cost as the current recomputation cost of the to-be-estimated tensor.

[0138] Figure 7 Another application scenario of memory allocation of the memory allocation and release decision system supporting dynamic recomputation according to the present disclosure is shown. As shown in FIG. 6, the memory pool 600 is used to store the tensors of the calculation logic nodes of the system. The memory pool 600 is divided into a plurality of memory spaces 601, 602, 603, 604, 605, 606, 607, 608, 609, and 610. The memory spaces 601, 602, 603, 604, 605, 606, 607, 608, 609, and 610 are respectively used to store the tensors of the calculation logic nodes 611, 612, 613, 614, 615, 616, 617, 618, 619, and 620. Figure 7The illustrated memory pool initialization component 120 obtains initial information of a data processing network or when a user initiates a recalculation function, applies for a continuous memory space of a predetermined size as a memory pool for the data processing network, and divides the memory pool of the predetermined size into at least two sub-memory pools. The data processing network is represented by tensors used by simplified logical computing nodes OP, such as tensors A, B, C, D, E, and F. Although six OP tensors are shown here, there are much more in the data processing network. The memory pool is evenly divided into five groups, i.e., sub-memory pool 1, sub-memory pool 2, sub-memory pool 3, sub-memory pool 4, and sub-memory pool 5. In actual use, m is usually used to represent the number of sub-memory pools.

[0139] In application, the memory space monitoring and allocation component 130 of the memory allocator 110 polls all the sub-memory pools from the next memory pool of the sub-memory pool from which the memory space allocated based on the previous request of the current request for allocation of a continuous memory space for data processing of a current computing logical node in the data processing network, to determine whether there is a continuous memory space greater than the requested size of the current request in one of all the sub-memory pools, and allocates the continuous memory space in the first polled sub-memory pool having the continuous memory space greater than the requested size of the current request to the computing logical node that makes the request. In the initial stage, each sub-memory pool has sufficient memory space, so the memory space is usually allocated in the sub-memory pools in a predetermined order according to the order number of the tensor (or computing logical node) of the applied memory space. In order to release the tensors in the above-mentioned manner, the tensors that are continuous in time and order are released at one time, resulting in a high recalculation cost when subsequent recalculation is performed. Therefore, in the initial stage, the continuous tensors are as far apart as possible in the memory space to which they belong, so that the tensors contained in the continuous (here, continuous refers to "continuous" in the memory address space) tensor subsequence obtained in the size-taking process are discontinuous in the running time order.

[0140] For this purpose, when there are only two sub-memory pools, as shown in the first schematic table in Figure 7 , the memory space monitoring and allocation component 130 alternately allocates the requested memory space in the two sub-memory pools according to the order of the current request. That is, the first sub-memory pool sequentially allocates memory from its starting address, and tensors A, C, and E are stored from the starting position of the sub-memory pool 1; the second sub-memory pool allocates memory in reverse order from its end address, and tensors B, D, and F are applied for memory space from the end address.

[0141] In the case that there are more than two sub-memory pools satisfying the requested memory space size, for example, n = 5, the memory space monitoring allocation component 130 allocates the memory space in the tensor order based on the order number of the current request, with the number of sub-memory pools n as the modulus, in the sub-memory pool numbered m, the remainder of the order number relative to the modulus n is also 1 + (m-1) * k, where k is an integer coprime with n, wherein the adjacent two sub-memory pools of the n sub-memory pools are allocated in sequence from the starting address in one case and in reverse order from the end address in the other case. As shown in the second schematic table of Figure 7 , which represents the remainder of the tensor number relative to the modulus 5 when k = 2, from m = 1-5, the remainders are respectively: 1, 3, 5, 7, 9, and since some remainders are greater than 5, the actual remainders are respectively: 1, 3, 0, 2, 4, and thus the storage order in the tensor A, B, C, D, E, F sub-memory pool is: tensor A, C, E, B, D, F, wherein tensor A and F are in the same sub-memory pool 1. In addition, in order to reduce the occurrence of fragmented memory blocks at the connection between adjacent two sub-memory pools, the starting points of adjacent sub-memory pools are just opposite when allocating memory space. As shown in Figure 7 , for example, sub-memory pool 1 starts to allocate memory space in sequence from the left starting point of its memory space, while sub-memory pool 2 starts to allocate memory space in reverse order from the right end point of its memory space, so that the empty space between sub-memory pool 1 and 2 is as large as possible, so as to reduce the opportunity of forming fragmented memory as much as possible. More importantly, since sub-memory pool 3 starts to allocate memory space in sequence from the left starting point of its memory space, there will be no fragmented memory space between sub-memory pool 2 and 3. Moreover, since tensor A and tensor B are consecutive tensors in time sequence, in order to avoid being released at the same time through the tensor sub-sequence mode, they are isolated in the address space through the above-mentioned memory allocation mode, so it is difficult for them to exist in the same sub-sequence at the same time, thereby avoiding the situation that subsequent recalculation cost increases due to being released at the same time.

[0142] As shown in the third schematic table of Figure 7 , which represents the remainder of the tensor number relative to the modulus 5 when k = 3, from m = 1-5, the remainders are respectively: 1, 4, 7, 10, 13, and since some remainders are greater than 5, the actual remainders are respectively: 1, 4, 2, 0, 3, and thus the storage order in the tensor A, B, C, D, E, F sub-memory pool is: tensor A, D, B, E, C, F, wherein tensor A and F are in the same sub-memory pool 1. In addition, in order to reduce the occurrence of fragmented memory blocks at the connection between adjacent two sub-memory pools, the starting points of adjacent sub-memory pools are just opposite when allocating memory space. As shown in Figure 7As shown, for example, the sub-memory pool 1 allocates memory space sequentially from the left start point of its memory space, while the sub-memory pool 2 allocates memory space in reverse order from the right end point of its memory space, so that the spare space between the sub-memory pool 1 and 2 is as large as possible, so as to reduce the opportunity of forming fragmented memory as much as possible. More importantly, since the sub-memory pool 3 allocates memory space sequentially from the left start point of its memory space, therefore, the space between the sub-memory pool 2 and 3 is completely free of fragmented memory. Moreover, since the tensor A and the tensor B are consecutive tensors in time sequence, in order to avoid being released simultaneously by the tensor sub-sequence mode, therefore, they are isolated in the address space by the above-mentioned memory allocation mode, so it is difficult for them to exist in the same sub-sequence simultaneously, thereby avoiding the situation of being released simultaneously and causing the subsequent recalculation cost to increase.

[0143] In the case of a large number of sub-memory pools n, in order to make the time-sequentially adjacent or close tensors possibly far apart in the storage space, k is the number closest to n / 2 and farther from 1. For example, in the case of n=5, k=5 / 2, which is the same as 2 and 3, so 3 is selected which is farther from 1. Of course, it is also feasible to select k=2.

[0144] It should be noted that when the memory space monitoring and allocation component 130 polls the current sub-memory pool in all sub-memory pools and confirms that the remaining memory space of the current sub-memory pool is less than the requested size of the current request for continuous memory space, the extended polling is extended to the adjacent sub-memory pool along the polled sub-memory pool, so that when the continuous memory space at the junction between the current sub-memory pool and its adjacent sub-memory pool is not less than the requested size of the current request for continuous memory space, the continuous space at the adjacent junction obtained by the extended polling is allocated to the computing logic node of the proposed request. As Figure 7As shown in the fourth schematic table, when the logical computing node of tensor N should allocate memory space in sub-memory pool 2 according to the grouped sub-memory pool, if the memory space required by the size of tensor N is 300M, but the remaining space on the left side of sub-memory pool 2 is 200M. However, there is still 100M remaining at the end of the right side of sub-memory pool 1 adjacent to sub-memory pool 2. Therefore, the memory space monitoring and allocation component 130 will continue to expand the query from the left side of sub-memory pool 2 to the left side of sub-memory pool 1, at this time, the expanded part of the memory space originally belonging to sub-memory pool 1 will be regarded as the memory space belonging to sub-memory pool 2. Thus, since the free memory space of the two adjacent parts is exactly 300M, which meets the needs of tensor N, the 300M of memory space will be allocated to tensor N. When the continuous memory space at the junction between the current sub-memory pool (for example, sub-memory pool 2) and its adjacent sub-memory pool (for example, sub-memory pool 1) is still less than the requested size of the continuous memory space of the current request, for example, less than the memory space required by tensor N, 300M, the next sub-memory pool (for example, sub-memory pool 3 and 4) is polled by skipping the current sub-memory pool and its adjacent sub-memory pool. If there is enough space in the next sub-memory pool, the memory space is allocated to tensor N.

[0145] By adopting the memory allocation and release decision system and method supporting dynamic recalculation according to the present disclosure, by dividing the video memory into a plurality of continuous sub-memory pools, i.e., into a plurality of continuous groups, and allowing the tensors in time sequence to be allocated to different sub-memory pools when there is enough available space in each sub-memory pool, the tensors in time sequence can be allocated to different sub-memory pools as much as possible, so that on a linear data processing network, the continuous video memory can be released as much as possible, which is actually releasing one tensor every several tensors, thereby reducing the release of a series of tensors, resulting in an increase in the recalculation cost caused by the release of memory in the later stage. More importantly, by selecting the sub-sequence of continuous tensors, the release of memory is no longer the release of one tensor with the minimum cost, but the release of which tensors can bring enough continuous video memory, which eliminates any unnecessary release (for example, in the case of releasing one tensor with the minimum cost, the space after the release may still not meet the space requirement, so the release must continue, and this repeated situation will lead to unnecessary release, i.e., the release of a tensor still cannot solve the OOM problem). Therefore, by releasing the sub-sequence of continuous tensors, only continuous video memory can be produced. Moreover, by regarding the idle video memory as a zero-cost tensor, the acquisition of the sub-sequence of continuous tensors will not cause an unselectable situation.

[0146] In summary, by employing the memory allocation and release decision system and method supporting dynamic recomputation according to this disclosure, and by dividing a pre-allocated memory pool of a predetermined threshold capacity into multiple sub-memory pools, the allocation of continuous tensors during network computation is reduced, minimizing their continuous distribution within the memory pool. Furthermore, by using a sliding window approach to obtain tensors within a larger continuous space for release, the high recomputation overhead caused by the release of continuous tensors during network computation is reduced. This also accelerates the acquisition of releasable continuous memory space and shortens network training time. Moreover, by adopting the method of releasing a predetermined size of continuous space as a whole, the memory release decision becomes more flexible and comprehensive, significantly reducing recomputation costs while achieving the memory release goal.

[0147] Furthermore, the inclusion of the cost of inverse recomputation as a factor in evaluating memory release in dynamic recomputation is primarily due to the existence of reversible modules in neural networks, where the input to a reversible module can be derived from its output. Examples include "flow-based models" such as NICE, RealNVP, and Glow. In these flow models, the Jacobian matrix of each reversible basic unit is a triangular matrix with a determinant of 1, making it convenient for solving the probability density distribution after the inverse transformation of a generative model. Additionally, reversible modules designed using Reversible Residual Networks (RevNet) can also be derived from the output (e.g., Figure 3 In the scenario shown, the inputs (e.g., x1, x2, and grad_x) are derived from y1 and y2 and the gradient (e.g., grad_y). RevNet replaces the residual module in ResNet with a reversible module, retaining only the output of the final layer. This significantly reduces GPU memory usage while achieving experimental results comparable to ResNet, with a computation time 1.5-2 times that of ResNet. Furthermore, common operations such as add_n, scalar_mul, and fully connected layers with reversible weights are all reversible. By incorporating reverse recomputation as a factor in memory release decisions, memory release decisions become more flexible and comprehensive, and the cost of recomputation is greatly reduced while achieving memory release goals.

[0148] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that those skilled in the art will understand that all or any step or component of the methods and apparatus of this disclosure can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software or a combination thereof. This is something that those skilled in the art can achieve by using their basic programming skills after reading the description of this disclosure.

[0149] Therefore, the object of the present disclosure can also be achieved by running a program or a set of programs on any computing device. The computing device can be a commonly known general-purpose device. Therefore, the object of the present disclosure can also be achieved only by providing a program product containing program code for implementing the method or device. That is, such a program product also constitutes the present disclosure, and a storage medium storing such a program product also constitutes the present disclosure. Obviously, the storage medium can be any commonly known storage medium or any storage medium developed in the future.

[0150] It is also necessary to point out that in the device and method of the present disclosure, obviously, the components or steps can be decomposed and / or recombined. These decompositions and / or recombination should be considered as equivalent solutions of the present disclosure. Moreover, the steps of performing the above series of processes can naturally be executed in time sequence according to the order of description, but do not necessarily have to be executed in time sequence. Some steps can be executed in parallel or independently of each other.

[0151] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A system for memory allocation and release decision supporting dynamic recomputation, comprising: a memory pool initialization component, which applies a continuous memory space of a predetermined size to a data processing network as a memory pool, and divides the memory pool of the predetermined size into at least two sub-memory pools; a memory space monitoring and allocating component, which is deployed in a memory allocator, polls all the sub-memory pools starting from a next memory pool of a sub-memory pool from which a memory space is allocated based on a previous request of a current request of a current computation logic node in the data processing network requesting allocation of a continuous memory space for its data processing, to determine whether there is a continuous memory space in one of the sub-memory pools larger than a requested size of the current request, and allocates a continuous memory space in a first polled sub-memory pool in which there is a continuous memory space larger than the requested size of the current request to the computation logic node making the request; a continuous tensor obtaining component, which sequentially looks for a continuous tensor sub-sequence in the memory pool in which a sum of occupied memory spaces is larger than a first size according to an address order of the memory pool by a Sieve method after the memory space monitoring and allocating component polls all the sub-memory pools once and determines that there is no continuous memory space in any of the sub-memory pools larger than the requested size of the current request; a recomputation cost evaluation component, which calculates a recomputation cost of each current tensor constituting each continuous tensor sub-sequence, and sums up to obtain a total recomputation cost of the continuous tensor sub-sequence; a memory releasing component, which sorts total recomputation costs of all continuous tensor sub-sequences according to sizes, and releases all tensors contained in a continuous tensor sub-sequence corresponding to a smallest total recomputation cost, thereby increasing a re-allocatable memory space of the memory pool. 2.The system for memory allocation and release decision supporting dynamic recomputation according to claim 1, wherein the memory space monitoring and allocating component alternately allocates requested memory spaces in two sub-memory pools in an order of current requests when there are only two sub-memory pools, wherein in a case that memory allocation is sequentially performed from a start address of a first sub-memory pool, memory allocation is performed in a reverse order from an end address of a second sub-memory pool. 3.The system for memory allocation and release decision supporting dynamic recomputation according to claim 1, wherein the memory space monitoring and allocating component, when there are more than two sub-memory pools satisfying a requested memory space size, allocates memory spaces in a tensor order with a remainder of a sequential number of a current request with respect to a modulus n of a number n of sub-memory pools being 1+(m-1)*k in a sub-memory pool numbered m, where k is an integer coprime with n, and wherein adjacent two of the n sub-memory pools are alternately allocated in a case that one is sequentially allocated from a start address and the other is allocated in a reverse order from an end address. 4.The system for memory allocation and release decision supporting dynamic recomputation according to claim 3, wherein k is a number closest to n / 2.

5. The memory allocation and deallocation decision system supporting dynamic recomputation according to any one of claims 1-4, wherein the memory space monitoring allocation component, upon polling a current sub-memory pool among all sub-memory pools and confirming that its remaining memory space is less than the requested size of the current request for contiguous memory space, extends the polling to an adjacent sub-memory pool along the polled sub-memory pool, so as to allocate contiguous memory space at the adjacent junction obtained by the extended polling to the computing logic node making the request, if the contiguous memory space at the adjacent junction between the current sub-memory pool and its adjacent sub-memory pool is not less than the requested size of the current request for contiguous memory space, and skips the polling of the next sub-memory pool from the current sub-memory pool and its adjacent sub-memory pool, if the contiguous memory space at the adjacent junction between the current sub-memory pool and its adjacent sub-memory pool is still less than the requested size of the current request for contiguous memory space.

6. The memory allocation and deallocation decision system supporting dynamic recomputation according to claim 5, wherein the contiguous tensor sub-sequence contains one, two or more than two tensors with their storage spaces contiguous to each other, or two or more than two tensors with free spaces between their storage spaces.

7. The memory allocation and deallocation decision system supporting dynamic recomputation according to claim 6, wherein the recomputation cost evaluation component evaluates the forward recomputation cost of each tensor to be evaluated among the tensors currently saved in the memory pool, wherein the free space between two tensors is regarded as a tensor with zero recomputation cost.

8. The memory allocation and deallocation decision system supporting dynamic recomputation according to claim 7, wherein the recomputation cost evaluation component comprises: a tensor usage history recording unit recording the tensor usage history and providing the time interval between the most recent usage time of the tensor to be evaluated and the current time; a tensor recursive query unit starting from each tensor to be evaluated of each contiguous tensor sub-sequence in the memory pool to recursively query its parent tensors and child tensors, and the recursive query terminates at the first parent tensor and the first child tensor of the tensor to be evaluated currently existing in the memory pool, thereby determining a first set of all parent tensors and all child tensors for evaluating the recomputation cost of the tensor to be evaluated, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the tensor to be evaluated, and all child tensors include all intermediate child tensors between the first child tensor and the tensor to be evaluated; a cost evaluation unit summing up the forward computation cost of each tensor to be evaluated of each contiguous tensor sub-sequence and the forward computation cost of each tensor in the first set, to obtain the forward computation cost of each tensor to be evaluated and the total recomputation cost of each contiguous tensor sub-sequence.

9. The memory allocation and deallocation decision system supporting dynamic recomputation according to claim 7, wherein the recomputation cost evaluation component comprises: ​ ​ ​ ​ The tensor usage history recording unit records tensor usage history and provides a time interval between a time point of a to-be-estimated tensor and a current time point; The tensor recursive query unit starts from each to-be-estimated tensor of each continuous tensor subsequence in the memory pool to recursively query its parent tensors and child tensors, and the recursive query terminates at the first parent tensor and the first child tensor of the to-be-estimated tensor which are currently present in the memory pool, thereby determining a first set of all parent tensors and all child tensors for estimating the re-computation cost of the to-be-estimated tensor, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the to-be-estimated tensor, and all child tensors include all intermediate child tensors between the first child tensor and the to-be-estimated tensor; The cost estimation unit sums the forward computation cost of each to-be-estimated tensor of each continuous tensor subsequence and the forward computation cost of each tensor in the first set to obtain the forward computation cost sum of each to-be-estimated tensor, calculates the re-computation cost time ratio between the forward computation cost sum of each to-be-estimated tensor and the time interval of the to-be-estimated tensor, sums the corresponding re-computation cost time ratios of all to-be-estimated tensors of each continuous tensor subsequence to obtain a re-computation cost time ratio sum, and calculates the ratio between the re-computation cost time ratio sum and the spatial size of each continuous tensor subsequence as the total re-computation cost of each continuous tensor subsequence.

10. A memory allocation and release decision method supporting dynamic re-computation, comprising: allocating, by a memory pool initialization component, a continuous memory space of a predetermined size as a memory pool for a data processing network, and dividing the memory pool of the predetermined size into at least two sub-memory pools; based on a current request of a current computation logic node in the data processing network requesting allocation of a continuous memory space for data processing thereof, polling, by a memory space monitoring allocation component deployed in a memory allocator, all sub-memory pools from a next memory pool of a sub-memory pool in which a memory space is allocated based on a previous request of the current request, to determine whether there is a continuous memory space greater than a requested size of the current request in one of all sub-memory pools, and allocating the continuous memory space in the first polled sub-memory pool in which there is a continuous memory space greater than the requested size of the current request to the computation logic node making the request; when it is determined by the memory space monitoring allocation component that there is no continuous memory space greater than the requested size of the current request in any sub-memory pool after polling all sub-memory pools once, sequentially searching, by a continuous tensor acquisition component, for continuous tensor subsequences in the memory pool in which the sum of occupied memory spaces is greater than a first size according to the address order of the memory pool by the Steiner method; calculating, by a re-computation cost evaluation component, the re-computation cost of each current tensor constituting each continuous tensor subsequence, and summing to obtain the total re-computation cost of the continuous tensor subsequence; The total weight computation cost of all continuous tensor sub-sequences is sorted by size through the memory release component, and the memory occupied by all tensors contained in the continuous tensor sub-sequence corresponding to the smallest total weight computation cost is released, thereby increasing the redistributable memory space of the memory pool.

11. The memory allocation and release decision method supporting dynamic re-computation according to claim 10, wherein the memory space monitoring allocation component, when there are only two sub-memory pools, alternately allocates the requested memory space in the two sub-memory pools in the order of current requests, wherein in the case of sequentially allocating memory from the start address of the first sub-memory pool, the second sub-memory pool allocates memory in reverse order from the end address.

12. The memory allocation and release decision method supporting dynamic re-computation according to claim 10, wherein the memory space monitoring allocation component, when there are more than two sub-memory pools that meet the requested memory space size, based on the order number of the current request, allocates memory space in the sub-memory pool numbered m with the same remainder 1+(m-1)*k as the order number relative to the modulus n, where k is an integer coprime with n, and the adjacent two of the n sub-memory pools allocate memory space in one from the start address in order and in the other from the end address in reverse order.

13. The memory allocation and release decision method supporting dynamic re-computation according to claim 12, wherein k is the number closest to n / 2.

14. The memory allocation and release decision method supporting dynamic re-computation according to any one of claims 9-13, wherein the memory space monitoring allocation component, when polling the current sub-memory pool in all sub-memory pools and confirming that the remaining memory space is less than the requested size of the current request, extends the polling to the adjacent sub-memory pool along the polled sub-memory pool, so that when the continuous memory space at the junction between the current sub-memory pool and its adjacent sub-memory pool is not less than the requested size of the current request, the continuous space at the adjacent junction obtained by the extended polling is allocated to the computing logic node making the request, and when the continuous memory space at the junction between the current sub-memory pool and its adjacent sub-memory pool is still less than the requested size of the current request, the polling of the next sub-memory pool is skipped.

15. The memory allocation and release decision method supporting dynamic re-computation according to claim 14, wherein the number of tensors contained in the continuous tensor sub-sequence is one, two or more than two tensors whose storage spaces are continuous, or two or more than two tensors with free space between their storage spaces.

16. The memory allocation and release decision method supporting dynamic re-computation according to claim 15, wherein The recompute cost evaluation component estimates the forward recompute cost of each tensor in the current saved tensors in the memory pool as the tensor to be evaluated, and considers the free space between two tensors as a tensor with zero recompute cost.

17. The method of claim 16, wherein the recompute cost evaluation by the recompute cost evaluation component comprises: recording tensor usage history and providing a time interval from the last usage time of a tensor to be evaluated to the current time by a tensor usage history recording unit; determining a first set of all parent tensors and all child tensors for estimating the recompute cost of the tensor to be evaluated by a tensor recursive query unit starting from each tensor to be evaluated in each continuous tensor sub-sequence in the memory pool recursively querying its parent tensors and child tensors, and the recursive query terminates at the first parent tensor and the first child tensor of the tensor to be evaluated that are currently existing in the memory pool, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the tensor to be evaluated, and all child tensors include all intermediate child tensors between the first child tensor and the tensor to be evaluated; summing the forward computation cost of each tensor to be evaluated in each continuous tensor sub-sequence and the forward computation cost of each tensor in the first set by a cost estimation unit to obtain the forward computation cost of each tensor to be evaluated and the total recompute cost of each continuous tensor sub-sequence.

18. The method of claim 16, wherein the recompute cost evaluation by the recompute cost evaluation component comprises: recording tensor usage history and providing a time interval from the last usage time of a tensor to be evaluated to the current time by a tensor usage history recording unit; determining a first set of all parent tensors and all child tensors for estimating the recompute cost of the tensor to be evaluated by a tensor recursive query unit starting from each tensor to be evaluated in each continuous tensor sub-sequence in the memory pool recursively querying its parent tensors and child tensors, and the recursive query terminates at the first parent tensor and the first child tensor of the tensor to be evaluated that are currently existing in the memory pool, wherein all parent tensors include all intermediate parent tensors between the first parent tensor and the tensor to be evaluated, and all child tensors include all intermediate child tensors between the first child tensor and the tensor to be evaluated; The forward calculation cost of each to-be-estimated tensor of each continuous tensor subsequence is summed by the cost estimation unit, and the forward calculation cost of each tensor in the first set is summed to obtain the forward calculation cost sum of each to-be-estimated tensor, and the re-calculation cost time ratio between the forward calculation cost sum of each to-be-estimated tensor and the time interval of the to-be-estimated tensor is calculated, the corresponding re-calculation cost time ratios of all to-be-estimated tensors of each continuous tensor subsequence are summed to obtain a re-calculation cost time ratio sum, and the ratio between the re-calculation cost time ratio sum and the spatial size of each continuous tensor subsequence is calculated as the total re-calculation cost of each continuous tensor subsequence.

Citation Information

Patent Citations

  • Memory management method and device, mobile terminal and storage medium

    CN109815162A

  • Tensor-based deep learning GPU memory management optimization method and system

    CN111078395A