Memory allocation method and device for graph and calculation integrated AI chip
By dynamically allocating the memory space of the AI chip, and optimizing memory allocation based on the topology structure of the computing graph and the life cycle of the output tensor, the memory allocation is solved, and the memory efficiency problem in the graph-computing integrated AI chip is achieved, achieving the maximum utilization of efficient computing and storage resources.
Patent Information
- Application Number
- CN202510467024.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, AI chips that integrate graphics and computing have problems with low efficiency and high energy consumption in memory allocation, resulting in increased computing resource usage and inability to effectively utilize the memory cell space.
By obtaining the topology of the calculation graph, the memory space of the AI chip is dynamically allocated according to the life cycle and type of the output tensor, and different preset allocation orders and specified rules are used to manage the output tensors with overlapping life cycles, including the use of storage resources such as local buffers, upper buffers and internal memory, and combined with greedy algorithms and adaptive heuristic algorithms to optimize memory allocation.
It improves memory allocation efficiency, reduces data transmission between storage units and computing units, reduces delay and energy consumption, and realizes efficient computing and maximizing the utilization of storage resources of AI chips.
Smart Images

Figure CN120371519A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a memory allocation method and device for an AI chip that integrates graphics and computing. Background Art
[0002] In the field of artificial intelligence, graph-computation integration mainly refers to the fusion optimization strategy of computational graph and operators. As a tool for depicting the architecture of neural networks in the field of deep learning, computational graph consists of a series of nodes connected by directed edges. Among them, nodes represent computing operations, and directed edges represent the paths of data flows. Operators are functions or operations embedded in these nodes and are responsible for performing specific computing tasks.
[0003] The core concept of graph-computation integration is to improve execution efficiency and reduce computing resource usage by subtly integrating and optimizing the combination mode of computational graphs and operators. To achieve this goal, the optimized computational graphs and operators need to be properly stored in limited storage units. This requires not only reasonable memory allocation for the tensors involved in each node operator in the neural network to maximize the space utilization of the storage unit, but also to ensure the scientific nature of address allocation so that the chip can quickly move and process the tensors required for calculation, thereby improving the chip's reasoning efficiency. Summary of the invention
[0004] In order to solve the problems in the related art, the embodiments of the present disclosure provide a memory allocation method and device for an AI chip that integrates graphics and computing. The present disclosure can effectively reduce the transmission of data between storage and computing units, thereby reducing latency and energy consumption.
[0005] In a first aspect, an embodiment of the present disclosure provides a memory allocation method for an AI chip that integrates graph and computation, wherein the AI chip implements graph and computation integration through a computation graph, wherein the computation graph includes a plurality of computing nodes, wherein the plurality of computing nodes are connected through directed edges, and each computing node has a corresponding output tensor, wherein the method includes:
[0006] Acquire a topological structure of the computation graph, wherein the topological structure includes an execution order relationship between the plurality of computation nodes;
[0007] According to the topological structure, obtaining the life cycle of the output tensor;
[0008] Obtain the type and required storage space of the output tensor;
[0009] Constructing data structure information of the output tensor according to the life cycle and required storage space of the output tensor;
[0010] When performing operations on the multiple computing nodes according to the described execution sequence relationship, at the start of the life cycle of the output tensor, allocate the available storage space of the AI chip for the output tensor, and at the termination of the life cycle of the output tensor, release the allocated available storage space;
[0011] When the multiple output tensors corresponding to the multiple computing nodes include output tensors with overlapping life cycles, for the output tensors with overlapping life cycles: perform a first pre-allocation of the available storage space of the AI chip for the output tensors with overlapping life cycles according to a specified rule in a first preset allocation order. After the first pre-allocation is successful, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles in the first preset allocation order; after the first pre-allocation fails, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping life cycles.
[0012] According to an embodiment of the present disclosure, obtaining the topological structure of the computation graph includes: performing a depth-first traversal or a breadth-first traversal on the multiple computing nodes, so as to obtain the topological structure of the computation graph.
[0013] According to an embodiment of the present disclosure, the life cycle of the output tensor starts at a first computing node and ends at a second computing node, where the first computing node is the computing node that outputs the output tensor, and the second computing node is the computing node that uses the output tensor as an input tensor for the last time.
[0014] According to an embodiment of the present disclosure, the types of the output tensors include any one of the following: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor;
[0015] Among them, the ordinary tensor is the output tensor other than the constant tensor, the exclusive tensor, the fixed tensor, and the non-exclusive tensor in the output tensors.
[0016] According to an embodiment of the present disclosure, the first preset allocation order is: constant tensor, fixed tensor, exclusive tensor, ordinary tensor, non-exclusive tensor;
[0017] The second preset allocation order is: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor;
[0018] The AI chip includes: a local buffer, an upper-layer buffer, and an internal memory.
[0019] According to an embodiment of the present disclosure, the specified rules include:
[0020] For the constant tensor, at the beginning of the life cycle, allocate the available storage space of the local buffer; at the end of the life cycle, release the allocated available storage space of the local buffer.
[0021] According to an embodiment of the present disclosure, the specified rules include:
[0022] For the fixed tensor, at the beginning of the life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the allocated available storage space of the upper-level buffer, the local buffer or the internal memory.
[0023] According to an embodiment of the present disclosure, obtaining the required storage space of the output tensor includes:
[0024] Set a temporary storage space for the exclusive tensor, the ordinary tensor, and the non-exclusive tensor respectively, and use the temporary storage space as the required storage space of the corresponding output tensor.
[0025] According to an embodiment of the present disclosure, the specified rules further include:
[0026] For the exclusive tensor, at the beginning of the life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space but greater than 0, reduce the required storage space and then allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is 0, allocate the available storage space of the local buffer; at the end of the life cycle, release the allocated available storage space of the upper-level buffer or the local buffer.
[0027] According to an embodiment of the present disclosure, the specified rules further include:
[0028] For the ordinary tensor, at the beginning of the life cycle, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the allocated available storage space of the local buffer or the internal memory.
[0029] According to an embodiment of the present disclosure, the specified rules further include:
[0030] When allocating the available storage space of the local buffer or the internal memory for the ordinary tensor, a two-dimensional memory model is constructed, where the abscissa of the two-dimensional memory model is the available storage space of the local buffer or the internal memory, and the ordinate is the time period;
[0031] According to the storage space size greedy algorithm and the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, and then select the algorithm that occupies less available storage space among the storage space size greedy algorithm and the adaptive heuristic algorithm to allocate the available storage space of the two-dimensional memory model.
[0032] According to an embodiment of the present disclosure, allocating the available storage space in the two-dimensional memory model according to the storage space size greedy algorithm includes:
[0033] Construct a corresponding rectangular data block according to the data structure information of the ordinary tensor, where the length and width of the corresponding rectangular data block are the required storage space and the life cycle of the ordinary tensor respectively;
[0034] Sort the rectangular data blocks of the ordinary tensor according to the size of the required storage space to obtain a sequence of rectangular data blocks;
[0035] Place the first rectangular data block in the sequence of rectangular data blocks in the two-dimensional memory model;
[0036] Iteratively perform the following operations until all the rectangular data blocks in the sequence of rectangular data blocks are placed in the two-dimensional memory model: traverse the unplaced rectangular data blocks from front to back, and select the rectangular data block with the smallest abscissa offset and place it in the two-dimensional memory model;
[0037] Wherein, the rectangular data blocks placed in the two-dimensional memory model do not overlap with each other.
[0038] According to an embodiment of the present disclosure, allocating the available storage space in the two-dimensional memory model according to the adaptive heuristic algorithm includes:
[0039] Construct a corresponding rectangular data block according to the data structure information of the ordinary tensor, where the length and width of the corresponding rectangular data block are the required storage space and the life cycle of the ordinary tensor respectively;
[0040] In the two-dimensional memory model, initialize a baseline, the abscissa offset of the baseline is 0, and the ordinate is the entire time period of the two-dimensional memory model;
[0041] Adaptively place one or more rectangular data blocks along the baseline until no other rectangular data blocks can be placed in the remaining time period of the baseline;
[0042] Iteratively perform the following operations until all the rectangular data blocks in the sequence of rectangular data blocks are placed in the two-dimensional memory model: Update the baseline such that the abscissa offset of the updated baseline is the available storage space with the smallest offset determined according to the placed rectangular data blocks, and the ordinate is the time period not occupied by the placed rectangular data blocks at the abscissa offset; adaptively place one or more rectangular data blocks along the updated baseline until the remaining segment length of the updated baseline cannot place other rectangular data blocks;
[0043] wherein the rectangular data blocks placed in the two-dimensional memory model do not overlap with each other.
[0044] According to an embodiment of the present disclosure, the specified rule further includes:
[0045] For the non-exclusive tensor, at the beginning of the life cycle, when the required storage space of the non-exclusive tensor is less than or equal to the maximum storage space of the corresponding computing node, reuse the maximum storage space of the corresponding computing node;
[0046] When the required storage space of the non-exclusive tensor is greater than the maximum storage space of the corresponding computing node, allocate the remaining storage space after the allocation of the constant tensor, the exclusive tensor, the fixed tensor, and the ordinary tensor is completed;
[0047] When the remaining storage space after the allocation of the constant tensor, the exclusive tensor, the fixed tensor, and the ordinary tensor is less than the required storage space of the non-exclusive tensor, allocate the remaining storage space after the allocation of the constant tensor, the exclusive tensor, and the fixed tensor, and then allocate the remaining storage space after the allocation of the non-exclusive tensor for the ordinary tensor;
[0048] At the end of the life cycle, release the allocated maximum storage space or the remaining storage space.
[0049] According to an embodiment of the present disclosure, each computing node also has a corresponding input tensor;
[0050] For each computing node, obtain the corresponding input tensor and output tensor, and use the maximum required storage space of the corresponding input tensor and output tensor as the maximum storage space of the computing node.
[0051] According to an embodiment of the present disclosure, the internal memory is a dual-rate synchronous dynamic random access memory.
[0052] Second aspect, embodiments of the present disclosure provide a memory allocation device for an AI chip for graph computing integration. The AI chip implements graph computing integration through a computation graph. The computation graph includes a plurality of computation nodes, and the plurality of computation nodes are connected by directed edges. Each computation node has a corresponding output tensor. The device includes:
[0053] A first acquisition module, configured to acquire the topological structure of the computation graph, where the topological structure includes the execution order relationship among the plurality of computation nodes;
[0054] A second acquisition module, configured to obtain the life cycle of the output tensor according to the topological structure;
[0055] A third acquisition module, configured to acquire the type and required storage space of the output tensor;
[0056] A construction module, configured to construct data structure information of the output tensor according to the life cycle and required storage space of the output tensor;
[0057] An allocation module, configured to, when performing operations on the plurality of computation nodes in accordance with the execution order relationship, at the start of the life cycle of the output tensor, allocate available storage space of the AI chip for the output tensor, and at the end of the life cycle of the output tensor, release the allocated available storage space;
[0058] A pre-allocation module, configured to, when the plurality of output tensors corresponding to the plurality of computation nodes include output tensors with overlapping life cycles, for the output tensors with overlapping life cycles: perform a first pre-allocation of the available storage space of the AI chip for the output tensors with overlapping life cycles according to a specified rule in a first preset allocation order, and after the first pre-allocation is successful, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles in the first preset allocation order; after the first pre-allocation fails, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping life cycles.
[0059] According to an embodiment of the present disclosure, the type of the output tensor includes any one of the following: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor, where the ordinary tensor is the output tensor other than the constant tensor, the exclusive tensor, the fixed tensor, and the non-exclusive tensor among the output tensors;
[0060] The first preset allocation order is: constant tensor, fixed tensor, exclusive tensor, ordinary tensor, non-exclusive tensor; the second preset allocation order is: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor;
[0061] The AI chip includes: a local buffer, an upper-layer buffer, and an internal memory.
[0062] According to an embodiment of the present disclosure, the specified rules include:
[0063] For the constant tensor, at the beginning of the life cycle, allocate the available storage space of the local buffer; at the end of the life cycle, release the allocated available storage space of the local buffer.
[0064] According to an embodiment of the present disclosure, the specified rules include:
[0065] For the fixed tensor, at the beginning of the life cycle, allocate the available storage space of the upper-layer buffer; when the available storage space of the upper-layer buffer is less than the required storage space, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the allocated available storage space of the upper-layer buffer, the local buffer, or the internal memory.
[0066] According to an embodiment of the present disclosure, the specified rules include:
[0067] For the exclusive tensor, at the beginning of the life cycle, allocate the available storage space of the upper-layer buffer; when the available storage space of the upper-layer buffer is less than the required storage space but greater than 0, reduce the required storage space and then allocate the available storage space of the upper-layer buffer; when the available storage space of the upper-layer buffer is 0, allocate the available storage space of the local buffer; at the end of the life cycle, release the allocated available storage space of the upper-layer buffer or the local buffer.
[0068] According to an embodiment of the present disclosure, the specified rules include:
[0069] For the ordinary tensor, at the beginning of the life cycle, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the allocated available storage space of the local buffer or the internal memory.
[0070] According to an embodiment of the present disclosure, the specified rules further include:
[0071] When allocating the available storage space of the local buffer or the internal memory for the ordinary tensor, a two-dimensional memory model is constructed, where the abscissa of the two-dimensional memory model is the available storage space of the local buffer or the internal memory, and the ordinate is the time period;
[0072] According to the storage space size greedy algorithm and the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, and then select the algorithm with less occupied available storage space among the storage space size greedy algorithm and the adaptive heuristic algorithm to allocate the available storage space of the two-dimensional memory model.
[0073] In a third aspect, an embodiment of the present disclosure provides an electronic device, including an AI chip, a memory, and a processor; wherein, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method according to any one of the first aspect.
[0074] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and characterized in that when the computer instructions are executed by a processor, the method according to any one of the first aspect is implemented.
[0075] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the method according to any one of the first aspect is implemented.
[0076] According to the technical solution provided by the embodiments of the present disclosure, by obtaining the topological structure of the computational graph, the topological structure includes the execution order relationship between the multiple computing nodes; according to the topological structure, obtaining the life cycle of the output tensor; obtaining the type and required storage space of the output tensor; constructing the data structure information of the output tensor according to the life cycle and required storage space of the output tensor; when performing the operations of the multiple computing nodes in accordance with the execution order relationship, at the start of the life cycle of the output tensor, allocating the available storage space of the AI chip for the output tensor, and at the end of the life cycle of the output tensor, releasing the allocated available storage space; when the multiple output tensors corresponding to the multiple computing nodes include output tensors with overlapping life cycles, for the output tensors with overlapping life cycles: performing a first pre-allocation of the available storage space of the AI chip for the output tensors with overlapping life cycles according to a specified rule in a first preset allocation order, and after the first pre-allocation is successful, allocating the available storage space of the AI chip for the output tensors with overlapping life cycles in the first preset allocation order; after the first pre-allocation fails, allocating the available storage space of the AI chip for the output tensors with overlapping life cycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping life cycles.
[0077] When the present disclosure meets the computing requirements of various output tensors in the computational graph, it can reduce the generation of memory fragmentation, improve the efficiency of memory allocation, and at the same time reduce the data transmission between the storage unit and the computing unit, thereby reducing the latency and energy consumption in the graph computing integrated architecture and achieving the optimal performance of the AI chip.
[0078] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In combination with the accompanying drawings, through the following detailed description of non-limiting embodiments, other features, objectives, and advantages of the present disclosure will become more obvious. In the drawings:
[0080] Figure 1 Show a flowchart of a method for memory allocation of an AI chip for graph computing integrated according to an embodiment of the present disclosure;
[0081] Figure 2 Show a schematic diagram of performing a depth-first traversal on multiple computing nodes in a computational graph in an embodiment of the present disclosure;
[0082] Figure 3Schematic diagram showing another implementation of the present disclosure for performing a depth - first traversal on multiple computing nodes in a computational graph;
[0083] Figure 4 Schematic diagram showing yet another implementation of the present disclosure for performing a depth - first traversal on multiple computing nodes in a computational graph;
[0084] Figure 5 Schematic diagram showing the execution time of an output tensor in a computational graph according to an embodiment of the present disclosure;
[0085] Figure 6 Schematic diagram showing the allocation of available storage space in the two - dimensional memory model according to the storage - space - size - based greedy algorithm in an embodiment of the present disclosure;
[0086] Figure 7 Schematic diagram showing the allocation of available storage space in the two - dimensional memory model according to the adaptive heuristic algorithm in an embodiment of the present disclosure;
[0087] Figure 8 Schematic block diagram showing a memory allocation device for a graph - computing - integrated AI chip according to an embodiment of the present disclosure. Detailed implementation manners
[0088] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for clarity, parts irrelevant to the description of the exemplary embodiments are omitted in the drawings.
[0089] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the existence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0090] In addition, it should be noted that, without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0091] In the present disclosure, if it involves operations of obtaining user information or user data or presenting user information or user data to others, such operations are all operations authorized, confirmed by the user, or actively selected by the user.
[0092] As mentioned above, the goal of graph - computing integration is to improve the execution efficiency and reduce the consumption of computing resources by optimizing the combination of computational graphs and operators. However, the optimized computational graphs and operators need to be stored in limited storage units. Therefore, how to reasonably allocate the memory for the tensors of each node in the computational graph has become a technical problem to be solved urgently.
[0093] The present disclosure provides a memory allocation method for an AI chip for graph computing integration. The AI chip implements graph computing integration through a computation graph. The computation graph includes a plurality of computation nodes, and the plurality of computation nodes are connected by directed edges. Each computation node has a corresponding output tensor. The method includes:
[0094] Obtain the topological structure of the computation graph, where the topological structure includes the execution order relationship between the plurality of computation nodes; according to the topological structure, obtain the life cycle of the output tensor; obtain the type and required storage space of the output tensor; construct the data structure information of the output tensor according to the life cycle and required storage space of the output tensor; when performing operations on the plurality of computation nodes in accordance with the execution order relationship, at the start of the life cycle of the output tensor, allocate the available storage space of the AI chip to the output tensor, and at the termination of the life cycle of the output tensor, release the allocated available storage space; when the plurality of output tensors corresponding to the plurality of computation nodes include output tensors with overlapping life cycles, for the output tensors with overlapping life cycles: perform a first pre-allocation of the available storage space of the AI chip for the output tensors with overlapping life cycles according to a specified rule in a first preset allocation order, and after the first pre-allocation is successful, allocate the available storage space of the AI chip to the output tensors with overlapping life cycles in the first preset allocation order; after the first pre-allocation fails, allocate the available storage space of the AI chip to the output tensors with overlapping life cycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping life cycles.
[0095] The present disclosure significantly enhances the ability of the AI chip in processing the computation graph, enabling the AI chip to quickly and accurately locate the required tensors. On this basis, the AI chip can efficiently transfer and compute these tensors, thereby greatly improving the efficiency of the inference process. At the same time, the present disclosure also optimizes the space utilization of the internal storage unit of the AI chip, achieving the maximum utilization of storage resources, not only accelerating the inference speed of the chip but also improving the space efficiency of the internal storage unit of the chip.
[0096] Figure 1 A flowchart showing a memory allocation method for an AI chip for graph computing integration according to an embodiment of the present disclosure. As Figure 1 shown, the memory allocation method includes the following steps S101 to S106:
[0097] In the present disclosure, the AI chip realizes graph computing integration through a computation graph. The computation graph includes multiple computation nodes, and the multiple computation nodes are connected by directed edges. Each computation node has a corresponding output tensor.
[0098] It is known that a computation graph is a directed acyclic graph (DAG). Computation nodes act as units for performing specific computation tasks (such as addition, multiplication, convolution, etc.) in the computation graph. Directed edges represent the data flow direction. In each computation node, specific computation tasks are implemented through operators (such as addition functions, multiplication functions, convolution functions, etc.). Each computation node has corresponding input tensors and an output tensor (including one or more input tensors and one output tensor), and is connected to other computation nodes through directed edges to form a complete computation process.
[0099] In step S101, obtain the topological structure of the computation graph, where the topological structure includes the execution order relationship among the multiple computation nodes.
[0100] Among them, perform a depth-first traversal or a breadth-first traversal on the multiple computation nodes to obtain the topological structure of the computation graph.
[0101] The following takes the depth-first traversal of the multiple computation nodes as an example for illustration.
[0102] According to an embodiment of the present disclosure, performing a depth-first traversal on the multiple computation nodes includes: accessing the vertices of the computation graph, taking the vertex as the current node, and iteratively performing the following operations until there are no unvisited adjacent computation nodes for the current node: accessing any adjacent computation node of the current node and taking the adjacent computation node as the new current node;
[0103] Backtracking to the previous node of the current node, taking the previous node as the new current node, and checking whether there are unvisited adjacent computation nodes for the current node;
[0104] If so, iteratively perform accessing any adjacent computation node of the current node and taking the adjacent computation node as the new current node; if not, iteratively perform backtracking to the previous node of the current node and taking the previous node as the new current node; until all the remaining computation nodes of the computation graph are visited.
[0105] Among them, the adjacent computation node of a computation node is a computation node that has a connected path with this computation node.
[0106] Figure 2 Shows a schematic diagram of performing a depth-first traversal on multiple computation nodes in an embodiment of the present disclosure.
[0107] In order to more clearly show the specific process of depth-first traversal, Figure 2 In the example shown, the branches in the computational graph are marked with different colors according to the execution order: red, blue, purple, and green. The specific execution process is: visit computational node 0, then use computational node 0 as the current node, visit computational node 1, then use computational node 1 as the new current node, visit computational node 3, then use computational node 3 as the new current node, visit computational node 6, then use computational node 6 as the new current computational node, visit computational node 7, then use computational node 7 as the new current computational node, and visit computational node 9. At this point, the red branch has been executed and traced back to the previous node of computational node 9 (computational node 7). Since computational node 7 does not have an unvisited adjacent computational node, it continues to trace back to the computational node 9. Compute node 6, and then visit another unvisited adjacent compute node of compute node 6 (compute node 8). The blue branch line is visited, and it traces back to the previous node of compute node 8 until it traces back to compute node 0. Then it visits another unvisited adjacent compute node of compute node 0 (compute node 2). Compute node 2 is used as the new current node, and it visits compute node 4. The purple branch line is executed, and then it traces back to compute node 2 and visits another unvisited adjacent compute node of compute node 2 (compute node 5). At this point, all the compute nodes in the computation graph have been visited, and the depth-first traversal ends.
[0108] Figure 3 Another schematic diagram of performing depth-first traversal on multiple computing nodes in a computing graph in an embodiment of the present disclosure is shown. Similarly, in Figure 3 In the example shown, the branches in the computational graph are marked with different colors according to the execution order: red, blue, purple, and green. Figure 3 The examples shown are similar to Figure 2 The difference of the example shown is that when computing node 2 is used as the new current node, computing node 5 of computing node 2 is first visited, and then backtracking to computing node 2 to visit computing node 4. That is, Figure 3 The examples shown are similar to Figure 2 The difference in the illustrated example is caused by the fact that computing node 2 has two adjacent computing nodes, and different adjacent computing nodes are selected.
[0109] Figure 4 FIG. 2 is a schematic diagram showing another embodiment of the present disclosure for performing a depth-first traversal on multiple computing nodes in a computing graph. Figure 4 As shown in the figure, the branches in the computational graph are marked with different colors according to the execution order: red, blue, purple, and green. Figure 4 The principle of performing depth-first traversal in the example shown is the same as Figure 2 and Figure 3The same as the illustrated example, which will not be elaborated here.
[0110] Performing a depth - first traversal on the computation graph is carried out along the branches of the computation graph. Since a computation graph may have multiple different branches, and a computation node may have multiple different adjacent computation nodes, the selected adjacent computation nodes will affect the final order of performing the depth - first traversal. Those skilled in the art should understand that the purpose of performing a depth - first traversal or a breadth - first traversal is to ensure that each computation node in the computation graph is visited to obtain the topological structure of the computation graph, and the selected traversal means and the specific traversal order do not affect the finally determined topological structure of the computation graph.
[0111] In step S102, according to the topological structure, obtain the lifecycle of the output tensor.
[0112] According to an embodiment of the present disclosure, the lifecycle of the output tensor starts from a first computation node and ends at a second computation node, where the first computation node is the computation node that outputs the output tensor, and the second computation node is the computation node that finally uses the output tensor as an input tensor.
[0113] In the present disclosure, the first computation node is also called the generating node, and the second computation node is also called the consuming node. When a computation node (i.e., the generating node) performs a computation and generates an output tensor, the lifecycle of this tensor officially starts. For example, in a neural network, the output tensor of the convolutional layer starts its lifecycle after its computation is completed. When this tensor is used as an input by the last computation node (i.e., the last consuming node), its lifecycle ends. For example, if a tensor is used by multiple nodes (such as in a fully - connected layer or an activation function), its lifecycle ends after the last node that uses it completes.
[0114] In step S103, obtain the type and required storage space of the output tensor.
[0115] According to an embodiment of the present disclosure, the types of the output tensor include any one of the following: constant tensor, fixed tensor, exclusive tensor, non - exclusive tensor, ordinary tensor; where the ordinary tensor is the output tensor other than the constant tensor, the exclusive tensor, the fixed tensor, and the non - exclusive tensor among the output tensors.
[0116] The definitions of each type of the output tensor are described in detail below:
[0117] Constant tensor (const): A tensor of const type is constant data and does not change during the computation process.
[0118] Fixed tensor: A tensor of the fixed type has a fixed size but may be read from or written to multiple times during computation.
[0119] Exclusive tensor: A tensor of the exclusive type requires exclusive access when being accessed, meaning that no other tensor should access the same storage space at the same time.
[0120] Ordinary tensor: A tensor of the tensor type is the most general type and has no specific access pattern or size limit during computation.
[0121] Inclusive tensor: A tensor of the inclusive type can reuse the available storage space in the previous nodes, if available, which means they may not require additional storage space.
[0122] In step S104, construct the data structure information of the output tensor according to the life cycle and required storage space of the output tensor.
[0123] That is, construct a data structure information (token) for each output tensor in the computation graph. The data structure information of each output tensor includes information about the life cycle of the tensor and the size of the required storage space, so as to provide reference information for allocating memory for the output tensor subsequently.
[0124] According to an embodiment of the present disclosure, respectively set up a temporary storage space for the exclusive tensor, the ordinary tensor, and the inclusive tensor, and use the temporary storage space as the required storage space for the corresponding output tensor.
[0125] For the three types of tensors, namely exclusive tensors, ordinary tensors, and inclusive tensors, since they do not have a fixed size and cannot directly obtain their required storage space like constant tensors and fixed tensors, therefore, a temporary storage space needs to be correspondingly set up for these three types of output tensors, and the temporary storage space is used as the required storage space for the corresponding tensors.
[0126] In step S105, when performing the operations of the multiple computing nodes in accordance with the execution order relationship, at the beginning of the life cycle of the output tensor, allocate the available storage space of the AI chip for the output tensor, and at the end of the life cycle of the output tensor, release the allocated available storage space.
[0127] In the present disclosure, when a tensor is no longer needed, the memory it occupies should be released, and the released memory will be returned to the memory pool. This mechanism reduces the overhead of memory allocation and improves the efficiency of memory usage.
[0128] As we know, the computational graph clarifies the dependencies between computational nodes. The computation of a computational node may depend on the output of other computational nodes. These dependencies determine the execution order of the computational nodes. The dependencies between computational nodes are represented by directed edges, which point from the dependent node to the dependent node in the computational graph. When executing a computational graph, it is usually executed from a computational node that does not have any pre-dependent dependencies (for example, Figures 2 - 4 Once all the pre-dependent nodes of a computation node have been executed and their output data are ready, the node can start executing. After a computation node is executed, its output will be passed along the directed edges to the subsequent dependent computation nodes, and these nodes will also be executed in turn after receiving the required data.
[0129] However, when operations of multiple computing nodes are executed according to the execution order relationship, the corresponding output tensors of each computing node may have overlapping life cycles.
[0130] Figure 5 Show Figures 2 - 4 Schematic diagram of the life cycle of output tensors in the computational graph shown.
[0131] like Figure 5 As shown, computing nodes 0 to 8 have corresponding output tensors 0 to 8, and finally the final result is obtained and output after the operation is performed at computing node 9. Among them, the life cycle of output tensor 0 starts at computing node 0 and ends at computing node 2, the life cycle of output tensor 1 starts at computing node 1 and ends at computing node 3, the life cycle of output tensor 2 starts at computing node 2 and ends at computing node 5, the life cycle of output tensor 3 starts at computing node 3 and ends at computing node 6, the life cycle of output tensor 4 starts at computing node 4 and ends at computing node 6, the life cycle of output tensor 5 starts at computing node 5 and ends at computing node 6, the life cycle of output tensor 6 starts at computing node 6 and ends at computing node 8, the life cycle of output tensor 7 starts at computing node 7 and ends at computing node 9, and the life cycle of output tensor 8 starts at computing node 8 and ends at computing node 9.
[0132] Since conflicts may occur when allocating memory for output tensors with overlapping lifecycles, the inventors classify the output tensors and manage the output tensors with overlapping lifecycles by specifying rules.
[0133] In step S106, when multiple output tensors corresponding to the multiple computing nodes include output tensors with overlapping lifecycles, for the output tensors with overlapping lifecycles: according to the first preset allocation order, the first pre-allocation of the available storage space of the AI chip is performed for the output tensors with overlapping lifecycles according to the specified rules. After the first pre-allocation is successful, the available storage space of the AI chip is allocated to the output tensors with overlapping lifecycles according to the first preset allocation order; after the first pre-allocation fails, the available storage space of the AI chip is allocated to the output tensors with overlapping lifecycles according to the second preset allocation order according to the specified rules, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping lifecycles.
[0134] The present disclosure enhances the ability of the AI chip in processing the computation graph, enabling the AI chip to efficiently transfer and compute different types of tensors in the computation graph, thereby improving the efficiency of the computation process, optimizing the utilization of the internal storage unit space of the AI chip, reducing the generation of fragmented resources, achieving the maximized utilization of storage resources, and reducing the computation latency and energy consumption.
[0135] The inventors noticed that since there are many types of output tensors, and the computation graph and tensors need to be stored in limited storage units, therefore, before actually performing memory allocation, pre-allocation (fake allocation) is first performed, that is, a simulated allocation process before formal allocation. In this way, on the one hand, the feasibility of the memory allocation scheme can be verified to ensure that all tensors can be successfully allocated, potential memory allocation problems can be discovered in advance, and adjustments can be made before actual allocation. For example, if it is found that some tensors cannot be allocated, immediate adjustments can be made; on the other hand, during the pre-allocation process, the impact of different allocation schemes on memory usage can be evaluated, and the optimal allocation scheme can be selected, which helps to reduce memory fragmentation and improve the allocation efficiency.
[0136] In the present disclosure, when determining output tensors with overlapping lifecycles, it is necessary to clarify the relationship of overlapping lifecycles between output tensors, which may include direct lifecycle overlap and extended lifecycle overlap. In the first case, if the lifecycles of two output tensors A and B have an intersection, that is, they both exist (are active) within the same time period, then output tensors A and B have direct lifecycle overlap. In the second case, if there is a set of output tensors and there is one or more output tensors such that any two output tensors in the set can be connected through a series of output tensors with direct lifecycle overlap (even if some output tensors do not directly overlap), then the output tensors in this set are said to form an extended lifecycle overlap; for example, output tensor A and output tensor B have overlapping lifecycles, output tensor B and output tensor C have overlapping lifecycles, and the lifecycles of output tensor A and output tensor C do not overlap. However, when considering output tensors A, B, and C as a whole set, because of the existence of output tensor B, it can be considered that they form a connected set of output tensors.
[0137] Furthermore, between two output tensors with overlapping lifecycles, it does not mean that the lifecycles of the two output tensors completely overlap, as long as there is an intersection between them. For example, if the lifecycle of output tensor A is from node 1 to node 3 and the lifecycle of output tensor B is from node 2 to node 5, then the lifecycles of output tensors A and B overlap.
[0138] According to an embodiment of the present disclosure, the first preset allocation order is: constant tensor, fixed tensor, exclusive tensor, ordinary tensor, non-exclusive tensor; the second preset allocation order is: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor.
[0139] According to an embodiment of the present disclosure, the AI chip includes: a local buffer, an upper buffer, and an internal memory. Among them, the internal memory is a dual-rate synchronous dynamic random access memory.
[0140] The following will further illustrate the specified rules according to the types of output tensors:
[0141] For the constant tensor, at the beginning of the lifecycle, allocate the available storage space of the local buffer; at the end of the lifecycle, release the allocated available storage space of the local buffer. Since the constant tensor is fixed and unchangeable, it is suitable to be placed close to the processing unit to reduce access latency.
[0142] For the fixed tensor, at the beginning of its life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the available storage space of the allocated upper-level buffer, local buffer or internal memory. Such an allocation strategy can reduce potential conflicts and improve access efficiency.
[0143] For the exclusive tensor, at the beginning of its life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space but greater than 0, reduce the required storage space and then allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is 0, allocate the available storage space of the local buffer; at the end of the life cycle, release the available storage space of the allocated upper-level buffer or local buffer. If the available storage space of the upper-level buffer is tight, the required storage space of the exclusive tensor can be appropriately reduced for adaptation.
[0144] For the ordinary tensor, at the beginning of its life cycle, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the available storage space of the allocated local buffer or internal memory. This can maximize the utilization of the chip memory while ensuring sufficient storage space.
[0145] Among them, when allocating the available storage space of the local buffer or internal memory for the ordinary tensor, a two-dimensional memory model is constructed, where the abscissa of the two-dimensional memory model is the available storage space of the local buffer or internal memory, and the ordinate is the time period; according to the storage space size greedy algorithm and the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, and then select the algorithm that occupies less available storage space among the storage space size greedy algorithm and the adaptive heuristic algorithm to allocate the available storage space of the two-dimensional memory model.
[0146] For the non-exclusive tensor, at the beginning of its life cycle, when the required storage space of the non-exclusive tensor is less than or equal to the maximum storage space of the corresponding computing node, reuse the maximum storage space of the corresponding computing node; when the required storage space of the non-exclusive tensor is greater than the maximum storage space of the corresponding computing node, allocate the remaining storage space after the allocation of the constant tensor, the exclusive tensor, the fixed tensor, and the ordinary tensor is completed; when the remaining storage space after the allocation of the constant tensor, the exclusive tensor, the fixed tensor, and the ordinary tensor is less than the required storage space of the non-exclusive tensor, allocate the remaining storage space after the allocation of the constant tensor, the exclusive tensor, and the fixed tensor, and then allocate the remaining storage space after the allocation of the non-exclusive tensor for the ordinary tensor; at the end of the life cycle, release the allocated maximum storage space or the remaining storage space. By reusing the previous storage space, the use of memory can be reduced and the memory utilization efficiency can be improved.
[0147] According to an embodiment of the present disclosure, allocating the available storage space in the two-dimensional memory model according to the storage space size greedy algorithm includes:
[0148] Construct a corresponding rectangular data block according to the data structure information of the ordinary tensor, where the length and width of the corresponding rectangular data block are the required storage space and the life cycle of the ordinary tensor, respectively;
[0149] Sort the rectangular data blocks of the ordinary tensor according to the size of the required storage space to obtain a sequence of rectangular data blocks;
[0150] Place the first rectangular data block in the sequence of rectangular data blocks in the two-dimensional memory model;
[0151] Iteratively perform the following operations until all the rectangular data blocks in the sequence of rectangular data blocks are placed in the two-dimensional memory model: Traverse the unplaced rectangular data blocks from front to back and select the rectangular data block with the smallest abscissa offset to place in the two-dimensional memory model;
[0152] Wherein, the rectangular data blocks placed in the two-dimensional memory model do not overlap each other.
[0153] According to an embodiment of the present disclosure, allocating the available storage space in the two-dimensional memory model according to the adaptive heuristic algorithm includes:
[0154] Construct a corresponding rectangular data block according to the data structure information of the ordinary tensor, where the length and width of the corresponding rectangular data block are the required storage space and the life cycle of the ordinary tensor, respectively;
[0155] In the two-dimensional memory model, initialize a baseline, where the abscissa offset of the baseline is 0 and the ordinate is the entire time period of the two-dimensional memory model;
[0156] Adaptively place one or more rectangular data blocks along the baseline until no other rectangular data blocks can be placed in the remaining time period of the baseline;
[0157] Iteratively perform the following operations until all the rectangular data blocks in the sequence of rectangular data blocks are placed in the two-dimensional memory model: update the baseline such that the abscissa offset of the updated baseline is the available storage space with the minimum offset determined according to the placed rectangular data blocks, and the ordinate is the time period not occupied by the placed rectangular data blocks at the abscissa offset; adaptively place one or more rectangular data blocks along the updated baseline until no other rectangular data blocks can be placed in the remaining segment length of the updated baseline;
[0158] Among them, the rectangular data blocks placed in the two-dimensional memory model do not overlap with each other.
[0159] Next, through two specific embodiments, the storage space size greedy algorithm and the adaptive heuristic algorithm will be described respectively. Figure 6 A schematic diagram showing the allocation of the available storage space in the two-dimensional memory model according to the storage space size greedy algorithm in the embodiments of the present disclosure. Figure 7 A schematic diagram showing the allocation of the available storage space in the two-dimensional memory model according to the adaptive heuristic algorithm in the embodiments of the present disclosure.
[0160] Suppose a computational graph includes ordinary tensors 0-7. It is known that the data structure information of each ordinary tensor includes: the required storage space and the life cycle. Then, the data structure information of ordinary tensors 0-7 is assumed to be:
[0161] Ordinary tensor 0: 32size, nodes 0-node 2; ordinary tensor 1: 28size, nodes 1-node 5; ordinary tensor 2: 36size, nodes 2-node 6; ordinary tensor 3: 16size, nodes 3-node 6; ordinary tensor 4: 8size, nodes 4-node 6; ordinary tensor 5: 64size, nodes 5-node 8; ordinary tensor 6: 10size, nodes 6-node 9; ordinary tensor 7: 40size, nodes 7-node 9.
[0162] Therefore, corresponding rectangular data blocks are constructed according to the data structure information of ordinary tensors, and rectangular data blocks block0-7 of ordinary tensors 0-7 are obtained, which are respectively: block0: 32 * [0-2], block1: 28 * [1-5], block2: 36 * [2-6], block3: 16 * [3-6], block4: 8 * [4-6], block5: 64 * [5-8], block6: 10 * [6-9], block7: 40 * [7-9]; A two-dimensional storage model is constructed, with the abscissa being the required storage space and the ordinate being the time period.
[0163] In Figure 6 the illustrated embodiment, the above rectangular data blocks block0-7 are sorted from largest to smallest according to the size of the required storage space, and a rectangular data block sequence [block5, block7, block2, block0, block1, block3, block6, block4] is obtained.
[0164] Then starting from the first rectangular data block block5, in the constructed two-dimensional memory model, traverse the unplaced rectangular data blocks in the rectangular data block sequence from front to back. Each time a placement is made, select the rectangular data block with the smallest abscissa offset for placement, and the rectangular data blocks cannot overlap with each other. The final placement result is as Figure 6 shown, where the numbers marked in the red circles represent the placement order of each rectangular data block.
[0165] In Figure 7 the illustrated embodiment, in the two-dimensional memory model, at the abscissa offset of 0 and the entire time period of the two-dimensional memory model for the ordinate, a baseline is initialized; block1 and block5 are placed along the initialized baseline. After placement, the remaining time periods in the initialized baseline are nodes 0-1 and nodes 8-9, and no other rectangular data blocks can be placed. The baseline needs to be updated. The abscissa offset of the updated baseline is the required storage space 28 of block1. At the position of the updated baseline, block0 is placed, and so on. By iteratively updating the baseline, the abscissa offset of the updated baseline is the available storage space with the smallest offset determined according to the placed rectangular data blocks, and the ordinate is the time period not occupied by the placed rectangular data block sequence at the abscissa offset. The final placement result is as Figure 7 shown, where the numbers marked in the red circles represent the placement order of each rectangular data block.
[0166] According to an embodiment of the present disclosure, each computing node also has a corresponding input tensor.
[0167] In the present disclosure, for each of the computing nodes, the corresponding input tensor and output tensor are obtained, and the maximum required storage space of the corresponding input tensor and output tensor is used as the maximum storage space of the computing node.
[0168] Figure 8 FIG. shows a structural block diagram of a memory allocation device for a graph computing integrated AI chip according to an embodiment of the present disclosure. Among them, the device can be implemented as part or all of an electronic device through software, hardware, or a combination of both.
[0169] As Figure 8 shown, the memory allocation device 800 includes a first acquisition module 810, a second acquisition module 820, a third acquisition module 830, an allocation module 840, and a pre-allocation module 850.
[0170] The first acquisition module 810 is configured to acquire the topological structure of the computation graph, where the topological structure includes the execution sequence relationship between the multiple computing nodes;
[0171] The second acquisition module 820 is configured to obtain the life cycle of the output tensor according to the topological structure;
[0172] The third acquisition module 830 is configured to acquire the type and required storage space of the output tensor;
[0173] The construction module is configured to construct the data structure information of the output tensor according to the life cycle and required storage space of the output tensor;
[0174] The allocation module 840 is configured to allocate the available storage space of the AI chip to the output tensor at the start of the life cycle of the output tensor when performing the operations of the multiple computing nodes in accordance with the execution sequence relationship, and release the allocated available storage space at the end of the life cycle of the output tensor;
[0175] The pre - allocation module 850 is configured to, when the multiple output tensors corresponding to the multiple computing nodes include output tensors with overlapping lifecycles, for the output tensors with overlapping lifecycles: perform a first pre - allocation of the available storage space of the AI chip for the output tensors with overlapping lifecycles according to a specified rule in a first preset allocation order; after the first pre - allocation is successful, allocate the available storage space of the AI chip for the output tensors with overlapping lifecycles in the first preset allocation order; after the first pre - allocation fails, allocate the available storage space of the AI chip for the output tensors with overlapping lifecycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping lifecycles.
[0176] According to an embodiment of the present disclosure, the types of the output tensors include any one of the following: constant tensor, fixed tensor, exclusive tensor, non - exclusive tensor, ordinary tensor, where the ordinary tensor is the output tensor other than the constant tensor, the exclusive tensor, the fixed tensor, and the non - exclusive tensor among the output tensors; the first preset allocation order is: constant tensor, fixed tensor, exclusive tensor, ordinary tensor, non - exclusive tensor; the second preset allocation order is: constant tensor, fixed tensor, exclusive tensor, non - exclusive tensor, ordinary tensor; the AI chip includes: a local buffer, an upper - layer buffer, and an internal memory.
[0177] According to an embodiment of the present disclosure, the specified rule includes: for the constant tensor, at the start of the lifecycle, allocate the available storage space of the local buffer; at the end of the lifecycle, release the allocated available storage space of the local buffer.
[0178] According to an embodiment of the present disclosure, the specified rule includes: for the fixed tensor, at the start of the lifecycle, allocate the available storage space of the upper - layer buffer; when the available storage space of the upper - layer buffer is less than the required storage space, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the lifecycle, release the allocated available storage space of the upper - layer buffer, the local buffer, or the internal memory.
[0179] According to an embodiment of the present disclosure, the specified rules include: for the exclusive tensor, at the beginning of the life cycle, allocate the available storage space of the upper buffer; when the available storage space of the upper buffer is less than the required storage space but greater than 0, reduce the required storage space and then allocate the available storage space of the upper buffer; when the available storage space of the upper buffer is 0, allocate the available storage space of the local buffer; at the end of the life cycle, release the allocated available storage space of the upper buffer or the local buffer.
[0180] According to an embodiment of the present disclosure, the specified rules include: for the ordinary tensor, at the beginning of the life cycle, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the allocated available storage space of the local buffer or the internal memory.
[0181] According to an embodiment of the present disclosure, the specified rules further include: for the ordinary tensor, when allocating the available storage space of the local buffer or the internal memory, construct a two-dimensional memory model, where the abscissa of the two-dimensional memory model is the available storage space of the local buffer or the internal memory, and the ordinate is the time period; according to the storage space size greedy algorithm and the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, and then select the algorithm that occupies less available storage space among the storage space size greedy algorithm and the adaptive heuristic algorithm to allocate the available storage space of the two-dimensional memory model.
[0182] The present disclosure can effectively reduce the transmission of data between the storage and computing units in the computational graph, thereby reducing latency and energy consumption and improving the processing efficiency of the AI chip.
[0183] The present disclosure also discloses an electronic device, which includes an AI chip, a memory, and a processor. Among them, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method according to the embodiment of the present disclosure.
[0184] In particular, according to an embodiment of the present disclosure, the method described above can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes program code for executing the above method. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium.
[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0186] The units or modules involved in the embodiments described in the present disclosure can be implemented in software or in programmable hardware. The described units or modules can also be provided in a processor, and the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0187] On the other hand, the present disclosure also provides a computer-readable storage medium, which can be the computer-readable storage medium included in the electronic device or computer system in the above embodiments; or it can exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the present disclosure.
[0188] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
Claims
1. A memory allocation method for an AI chip for graph computing integration, characterized in that, The AI chip implements graph computing integration through a computation graph. The computation graph includes multiple computing nodes, which are connected by directed edges. Each computing node has a corresponding output tensor. The method includes: Obtain the topological structure of the computation graph, where the topological structure includes the execution order relationship among the multiple computing nodes; Obtain the life cycle of the output tensor according to the topological structure; Obtain the type and required storage space of the output tensor; Construct the data structure information of the output tensor according to the life cycle and required storage space of the output tensor; When performing the operations of the multiple computing nodes according to the execution order relationship, at the start of the life cycle of the output tensor, allocate the available storage space of the AI chip for the output tensor, and at the end of the life cycle of the output tensor, release the allocated available storage space; When the multiple output tensors corresponding to the multiple computing nodes include output tensors with overlapping life cycles, for the output tensors with overlapping life cycles: perform the first pre-allocation of the available storage space of the AI chip for the output tensors with overlapping life cycles according to a specified rule in a first preset allocation order. After the first pre-allocation is successful, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles in the first preset allocation order; after the first pre-allocation fails, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping life cycles.
2. The method according to claim 1, wherein The obtaining of the topological structure of the computation graph includes: performing a depth-first traversal or a breadth-first traversal on the multiple computing nodes to obtain the topological structure of the computation graph.
3. The method according to claim 1, wherein The life cycle of the output tensor starts from a first computing node and ends at a second computing node, where the first computing node is the computing node that outputs the output tensor, and the second computing node is the computing node that last uses the output tensor as an input tensor.
4. The method according to claim 1, characterized in that, The types of the output tensor include any one of the following: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor; Among them, the ordinary tensor is the output tensor other than the constant tensor, the exclusive tensor, the fixed tensor, and the non-exclusive tensor among the output tensors.
5. The method according to claim 4, wherein: The first preset allocation order is: constant tensor, fixed tensor, exclusive tensor, ordinary tensor, non-exclusive tensor; The second preset allocation order is: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor; The AI chip includes: a local buffer, an upper-layer buffer, and an internal memory.
6. The method according to claim 5, wherein The specified rule includes: For the constant tensor, at the beginning of its life cycle, allocate the available storage space of the local buffer; at the end of its life cycle, release the allocated available storage space of the local buffer.
7. The method according to claim 5, characterized in that, The specified rules include: For the fixed tensor, at the beginning of its life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of its life cycle, release the allocated available storage space of the upper-level buffer, the local buffer or the internal memory.
8. The method according to claim 5, characterized in that Obtaining the required storage space of the output tensor includes: Separate temporary storage spaces are set for the exclusive tensor, the ordinary tensor, and the non-exclusive tensor respectively, and the temporary storage spaces are used as the required storage spaces of the corresponding output tensors.
9. The method according to claim 8, wherein The specified rules also include: For the exclusive tensor, at the beginning of its life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space but greater than 0, reduce the required storage space and then allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is 0, allocate the available storage space of the local buffer; at the end of its life cycle, release the allocated available storage space of the upper-level buffer or the local buffer.
10. The method according to claim 8, characterized in that, The specified rules also include: For the ordinary tensor, at the beginning of its life cycle, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of its life cycle, release the allocated available storage space of the local buffer or the internal memory.
11. The method according to claim 10, wherein The specified rules also include: When allocating the available storage space of the local buffer or the internal memory for the ordinary tensor, construct a two-dimensional memory model, where the abscissa of the two-dimensional memory model is the available storage space of the local buffer or the internal memory, and the ordinate is the time period; According to the storage space size greedy algorithm and the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, and then select the algorithm that occupies less available storage space between the storage space size greedy algorithm and the adaptive heuristic algorithm to allocate the available storage space of the two-dimensional memory model.
12. The method according to claim 11, wherein Allocating the available storage space in the two-dimensional memory model according to the storage space size greedy algorithm includes: Construct a corresponding rectangular data block according to the data structure information of the ordinary tensor, where the length and width of the corresponding rectangular data block are the required storage space and the life cycle of the ordinary tensor respectively; Sort the rectangular data blocks of the ordinary tensor according to the size of the required storage space to obtain a sequence of rectangular data blocks; Place the first rectangular data block in the sequence of rectangular data blocks in the two-dimensional memory model; Iteratively perform the following operations until all the rectangular data blocks in the sequence of rectangular data blocks are placed in the two-dimensional memory model: Traverse the unplaced rectangular data blocks from front to back, and select the rectangular data block with the smallest abscissa offset and place it in the two-dimensional memory model; Among them, the rectangular data blocks placed in the two-dimensional memory model do not overlap with each other.
13. The method according to claim 11, wherein According to the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, including: Construct corresponding rectangular data blocks according to the data structure information of the ordinary tensor, where the length and width of the corresponding rectangular data blocks are the required storage space and the life cycle of the ordinary tensor respectively; In the two-dimensional memory model, initialize a baseline, the abscissa offset of the baseline is 0, and the ordinate is the entire time period of the two-dimensional memory model; Adaptively place one or more rectangular data blocks along the baseline until no other rectangular data blocks can be placed in the remaining time period of the baseline; Iteratively perform the following operations until all the rectangular data blocks in the sequence of rectangular data blocks are placed in the two-dimensional memory model: Update the baseline so that the abscissa offset of the updated baseline is the available storage space with the smallest offset determined according to the placed rectangular data blocks, and the ordinate is the time period not occupied by the placed rectangular data blocks at the abscissa offset; Adaptively place one or more rectangular data blocks along the updated baseline until no other rectangular data blocks can be placed in the remaining segment length of the updated baseline; Among them, the rectangular data blocks placed in the two-dimensional memory model do not overlap with each other.
14. The method according to claim 11, wherein The specified rules also include: For the non-exclusive tensor, at the beginning of the life cycle, when the required storage space of the non-exclusive tensor is less than or equal to the maximum storage space of the corresponding computing node, reuse the maximum storage space of the corresponding computing node; When the required storage space of the non-exclusive tensor is greater than the maximum storage space of the corresponding computing node, allocate the remaining storage space after the allocation of the constant tensor, the exclusive tensor, the fixed tensor, and the ordinary tensor is completed; When the remaining storage space after the allocation of the constant tensor, the exclusive tensor, the fixed tensor, and the ordinary tensor is less than the required storage space of the non-exclusive tensor, allocate the remaining storage space after the allocation of the constant tensor, the exclusive tensor, and the fixed tensor, and then allocate the remaining storage space after the allocation of the non-exclusive tensor for the ordinary tensor; At the end of the life cycle, release the allocated maximum storage space or the remaining storage space.
15. The method according to claim 14, wherein Each computing node also has a corresponding input tensor; For each computing node, obtain the corresponding input tensor and output tensor, and use the maximum required storage space of the corresponding input tensor and output tensor as the maximum storage space of the computing node.
16. The method according to claim 5, characterized in that, The internal memory is a dual-rate synchronous dynamic random access memory.
17. A memory allocation device for an AI chip for integrated graphics and computing, characterized in that, The AI chip realizes graph computing integration through a computation graph. The computation graph includes a plurality of computing nodes, and the plurality of computing nodes are connected by directed edges. Each computing node has a corresponding output tensor. The device includes: A first acquisition module configured to acquire the topological structure of the computation graph, where the topological structure includes the execution order relationship among the plurality of computing nodes; A second acquisition module configured to obtain the life cycle of the output tensor according to the topological structure; A third acquisition module configured to acquire the type and required storage space of the output tensor; A construction module configured to construct the data structure information of the output tensor according to the life cycle and required storage space of the output tensor; An allocation module configured to, when performing operations on the plurality of computing nodes in accordance with the execution order relationship, at the start of the life cycle of the output tensor, allocate the available storage space of the AI chip for the output tensor, and at the end of the life cycle of the output tensor, release the allocated available storage space; A pre-allocation module configured to, when the plurality of output tensors corresponding to the plurality of computing nodes include output tensors with overlapping life cycles, for the output tensors with overlapping life cycles: perform a first pre-allocation of the available storage space of the AI chip for the output tensors with overlapping life cycles according to a specified rule in a first preset allocation order, and after the first pre-allocation is successful, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles in the first preset allocation order; after the first pre-allocation fails, allocate the available storage space of the AI chip for the output tensors with overlapping life cycles according to the specified rule in a second preset allocation order, where the first preset allocation order and the second preset allocation order include different allocation orders set according to the types of the output tensors with overlapping life cycles.
18. The device according to claim 17, wherein: The type of the output tensor includes any one of the following: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor, where the ordinary tensor is the output tensor other than the constant tensor, the exclusive tensor, the fixed tensor, and the non-exclusive tensor among the output tensors; The first preset allocation order is: constant tensor, fixed tensor, exclusive tensor, ordinary tensor, non-exclusive tensor; the second preset allocation order is: constant tensor, fixed tensor, exclusive tensor, non-exclusive tensor, ordinary tensor; The AI chip includes: a local buffer, an upper buffer, and an internal memory.
19. The device according to claim 18, wherein The specified rule includes: For the constant tensor, at the start of the life cycle, allocate the available storage space of the local buffer; at the end of the life cycle, release the allocated available storage space of the local buffer.
20. The device according to claim 18, wherein The specified rule includes: For the fixed tensor, at the beginning of the life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the available storage space of the allocated upper-level buffer, local buffer or internal memory.
21. The device according to claim 18, characterized in that, The specified rules include: For the exclusive tensor, at the beginning of the life cycle, allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is less than the required storage space but greater than 0, reduce the required storage space and then allocate the available storage space of the upper-level buffer; when the available storage space of the upper-level buffer is 0, allocate the available storage space of the local buffer; at the end of the life cycle, release the available storage space of the allocated upper-level buffer or local buffer.
22. The device according to claim 18, characterized in that, The specified rules include: For the ordinary tensor, at the beginning of the life cycle, allocate the available storage space of the local buffer; when the available storage space of the local buffer is less than the required storage space, allocate the available storage space of the internal memory; at the end of the life cycle, release the available storage space of the allocated local buffer or internal memory.
23. The device according to claim 22, characterized in that, The specified rules further include: When allocating the available storage space of the local buffer or the internal memory for the ordinary tensor, construct a two-dimensional memory model, where the abscissa of the two-dimensional memory model is the available storage space of the local buffer or the internal memory, and the ordinate is the time period; According to the storage space size greedy algorithm and the adaptive heuristic algorithm, allocate the available storage space in the two-dimensional memory model, and then select the algorithm that occupies less available storage space among the storage space size greedy algorithm and the adaptive heuristic algorithm to allocate the available storage space of the two-dimensional memory model.
24. An electronic device, characterized in that, It includes an AI chip, a memory and a processor; wherein, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 16.
25. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instruction is executed by the processor, it implements the method according to any one of claims 1 to 16.
26. A computer program product, including computer instructions, which implement the method according to any one of claims 1 to 16 when the computer instructions are executed by the processor.