A Compilation Optimization Method, Device, Equipment and Medium for Computational Graphs
By dynamically planning the computing graph to obtain the best slicing fusion scheme and converting the fusion group into a loop semantic node, the problem of frequent data handling between high-level and low-level storage layers is solved, and data processing efficiency and computing performance are improved.
Patent Information
- Application Number
- CN202510300601.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In artificial intelligence chips, when computing graphs perform ultra-large-scale tensor data processing, they need to frequently handle tensor data between the advanced storage layer and the low-level storage layer, resulting in a reduced data processing efficiency and affecting computing performance.
By using dynamic programming search on the computing graph, the best slicing and fusion scheme is searched, the search calculation is simplified using historical search results to the maximum extent, and the complex fusion groups are converted into simple loop semantic nodes, avoiding the transmission of tensor data between the advanced storage layer and the low-level storage layer.
The data processing efficiency and overall computing performance of the computing graph are improved, and the computing performance is significantly improved by reducing the number of data transmissions and optimizing the computing graph structure.
Smart Images

Figure CN119806542B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to computer technologies, and in particular, to a method, apparatus, device, and medium for compiling and optimizing a computation graph. Background Art
[0002] In an artificial intelligence chip, adopting a multi-level memory design is an effective means to balance cost and performance. Its characteristic is that the lower-level memory closer to the computing unit has faster data access and smaller capacity, while the higher-level memory farther from the computing unit has slower data access and larger capacity. A computation graph is a data structure representing the computing process of a neural network, composed of nodes and edges. Nodes represent operations, such as logical calculations or data transfers, and logical calculations or data transfers are ultimately executed by operators. Specifically, an operator is a data stream operation that moves the tensors to be computed from the higher-level memory to the lower-level memory, a mathematical calculation operation executed according to a specific logic, and a data stream operation that moves the calculation result from the lower-level memory back to the higher-level memory.
[0003] However, when the computation graph processes ultra-large-scale tensor data, it is necessary to frequently move tensor data between the high-level memory layer and the low-level memory layer, and increase the number of times of tensor data transmission between the memory and the computing unit as well as data read and write operations. This will significantly reduce the data processing efficiency of the computation graph, thereby affecting the computing performance of the computation graph. Summary of the Invention
[0004] Embodiments of the present invention provide a method for compiling and optimizing a computation graph to improve the data processing efficiency and overall computing performance of the computation graph.
[0005] In a first aspect, embodiments of the present invention provide a method for compiling and optimizing a computation graph, including: dividing the computation graph running on a multi-level memory chip to obtain multiple search regions;
[0006] Performing a search on each of the search regions using a dynamic programming search method to obtain an optimal splitting and fusion scheme;
[0007] When it is determined that the optimal splitting and fusion scheme includes a fusion group, converting the fusion group into a loop semantic node, where all operator nodes in the fusion group perform tensor transfer in the shared memory layer of the multi-level memory chip;
[0008] Modifying the computation graph according to the loop semantic nodes obtained from each of the search regions.
[0009] In a second aspect, embodiments of the present invention provide a device for compiling and optimizing a computation graph, including: a search region division module, configured to divide the computation graph running on a multi-level memory chip to obtain multiple search regions;
[0010] The optimal segmentation and fusion scheme search module is used to search for the optimal segmentation and fusion scheme for each of the search regions by using a dynamic programming search method;
[0011] The loop semantic node conversion module is used to convert the fusion group into a loop semantic node when it is determined that the optimal segmentation and fusion scheme includes a fusion group, where all operator nodes in the fusion group perform tensor transfer in the shared storage layer of the multi-level storage chip;
[0012] The computation graph modification module is used to modify the computation graph according to the loop semantic nodes obtained in each of the search regions.
[0013] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method described above when executing the program.
[0014] In a fourth aspect, an embodiment of the present invention provides a storage medium storing computer-executable instructions, on which a computer program is stored, wherein the program implements the method described above when executed by a processor.
[0015] The present invention searches for the optimal segmentation and fusion scheme for the computation graph by using a dynamic programming search method, thereby maximizing the use of historical search results to simplify the computational amount in the search process, and the operator nodes in the same fusion group in the optimal fusion scheme only perform operator calculations in the shared storage layer, thereby avoiding the transmission of tensor data between the high-level storage layer and the low-level storage layer, improving the data processing efficiency of the computation graph, and further improving the computational performance of the computation graph by converting the structurally complex fusion group into a simple loop semantic node. Description of the Drawings
[0016] Figure 1 is a flowchart of a compilation optimization method for a computation graph provided in Embodiment 1 of the present invention;
[0017] Figure 2 is a schematic structural diagram of a multi-level storage chip provided in Embodiment 1 of the present invention;
[0018] Figure 3 is a schematic diagram of the conversion process of a loop semantic node provided in Embodiment 1 of the present invention;
[0019] Figure 4 is a flowchart of a compilation optimization method for a computation graph provided in Embodiment 2 of the present invention;
[0020] Figure 5 is a schematic structural diagram of a compilation optimization device for a computation graph provided in Embodiment 3 of the present invention;
[0021] Figure 6 It is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. Detailed implementation manners
[0022] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that, for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0023] Embodiment 1
[0024] Figure 1 It is a flowchart of a compilation optimization method for a computation graph provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of compiling and optimizing a computation graph. This method can be executed by a compilation optimization device for the computation graph, and this device can be implemented in the form of hardware and / or software. As Figure 1 shown, this method includes:
[0025] Step S101, divide the computation graph running on a multi-level storage chip to obtain multiple search regions.
[0026] Optionally, dividing the computation graph running on a multi-level storage chip to obtain multiple search regions includes: identifying each operator node in the computation graph to obtain specified operator nodes that do not support splitting, where the specified operator nodes include global normalization operator nodes and user-defined operator nodes; filtering the specified operator nodes from the computation graph to obtain a processed computation graph; dividing the processed computation graph into regions to obtain multiple search regions, where each search region includes multiple connected operator nodes.
[0027] Specifically, the computation graph in this embodiment includes multiple operator nodes. The operator nodes represent operation logic calculation operations such as add and mul, or data transfer operations such as slice. The computation graph is represented in the form of Intermediate Representation (IR). And the computation graph in this embodiment runs on, for example Figure 2In the multi-level storage chip shown, where the L3 storage is the global storage space, which is the highest-level storage layer with the largest capacity but the slowest speed; the L2 storage is the shared storage layer that can be accessed by multiple computing units, such as SIPs, with a smaller capacity but a higher speed; the L1 storage is the exclusive storage of the computing unit, with the smallest capacity but the fastest speed. In this embodiment, compilation optimization is mainly performed for the computational graph that processes ultra-large-scale tensors, and the tensor data is the data processed by the operator nodes. The tensor data can have multiple dimensions, for example, 64x384x768, 64x3x224x224, and 64x224x224x3, etc. The highest dimension of the tensor represents the batch size (bs). For example, the highest dimension of each tensor data in the above examples is 64. Of course, this is only an example in this embodiment and does not limit the batch size of the ultra-large-scale tensors processed by the computational graph.
[0028] Among them, in this embodiment, the split-fusion operation is mainly performed in the shared storage layer, and the fusion mainly refers to IO fusion. When IO fusion is not performed, the output data of the operator nodes in the computational graph are all default stored in the L3 layer, that is, the highest-level storage layer. When the operator calculates, it loads data from L3 to L2, and then each computing unit SIP loads a part of the data to L1 for calculation. After the calculation is completed, it is written back to L3 from L1 through L2. This operator node is then considered to have been executed, and then the next operator node starts to load data from L3. Obviously, this method will increase the number of data transmissions between different storage levels. When IO fusion is adopted, the output of the previous operator node will not be written back to L3 but will be temporarily stored in L2. Then the next operator node can directly load data from L2, that is, it is only shown that the previous operator node loads data from L3, and the next operator node writes the calculation result after continuously executing the logic of these two operators back to L3. At this time, these two operator nodes are considered to be fused into one operator node. However, this requires that the output of the previous operator can be completely temporarily stored in L2. If the tensor scale is too large to be temporarily stored in L2, splitting is required.
[0029] It should be noted that in this embodiment, before performing the split-fusion, it is necessary to first identify each operator node in the computational graph to obtain the specified operator nodes that do not support splitting and filter them. For example, the global normalization operator node and the user-defined operator node, etc. Of course, this is only an example in this embodiment and does not limit the specific types of the specified operator nodes. And the operator nodes to be fused need to have a connection relationship. Therefore, in this embodiment, the filtered computational graph will be divided into regions to obtain multiple search regions, and each search region includes multiple connected operator nodes. By dividing the computational graph into regions, the ineffective fusion judgment of unconnected operator nodes can be avoided.
[0030] Step S102, for each search area, a dynamic programming search method is adopted to search for the optimal segmentation and fusion scheme.
[0031] Optionally, for each search area, a dynamic programming search method is adopted to search for the optimal segmentation and fusion scheme, including: for each search area, a dynamic programming search method is adopted to search for all the segmentation and fusion schemes for the operator nodes; calculate the segmentation benefits of each segmentation and fusion scheme in each search area, and take the segmentation and fusion scheme with the largest segmentation benefit as the optimal segmentation and fusion scheme in each search area.
[0032] Specifically, in this embodiment, when searching for the optimal segmentation and fusion scheme for each search area, a dynamic programming search method is specifically adopted to search for all the segmentation and fusion schemes for the operator nodes, and the segmentation and fusion scheme with the largest segmentation benefit is taken as the optimal segmentation and fusion scheme in each search area. And there will be a segmentation benefit only when there is a fusion group in the segmentation and fusion scheme. There is no segmentation benefit for the segmentation scheme without a fusion group. Because for the operator nodes in the same fusion group, the input of the previous operator node is not written back to L3, but is temporarily stored in L2 and used as the input for the subsequent operator node until the last operator node is written back to L3. Then these operator nodes can be called a fusion group, and the same segmentation method is used inside the fusion group to meet the requirement that the output of the previous operator node serves as the input of the subsequent operator node. Since the computational graph is represented in the form of IR, the representation method needs to be modified on the IR. Use fusion to enclose the IRs corresponding to multiple operator nodes that can be fused. For example, two IRs are wrapped to form a fused group IR, which contains two operators inside. That is, in terms of form: No.1 add -> No.2 mul becomes No.1 fusion(add+mul), and two IRs become one IR, that is, two separate operator nodes become a fusion group. Of course, in this embodiment, only the fusion of two operator nodes is taken as an example, and the fusion method for multiple operator nodes is roughly the same, which will not be elaborated in this embodiment.
[0033] Among them, the purpose of splitting in this embodiment is to cache the output tensor after splitting the previous operator on the L2 as the input for the next operator node, for the purpose of performing IO fusion. Splitting and fusion are closely related. The dynamic search method is to search out which operator nodes are fused into a fusion group. In order for the processed tensor data to be cached by the L2, it is necessary to split which dimensions to minimize the data transfer time of the computational graph, that is, to maximize the benefit. After splitting, the shape becomes smaller. For example, 6x3x224x224 can be split into 3 2x3x224x224. This also means that the operator node needs to be processed three times, that is, the operator node changes from one to three to process all this data. The tensor data processed each time becomes bs2. Therefore, when splitting, specifically, bs2 is split from the original L3 bs6 tensor data each time, transported to the L2 to call the operator for calculation. After the three splits and calculations are completed, the results of bs2 on the three L2s still need to be transported back to the L3 and concatenated into L3 bs6 as the input for the next operator node.
[0034] In a specific implementation, in the dynamic programming scheme search of this embodiment, when calculating the benefit during the search process, it is assumed that some operator nodes perform L2 IO fusion for calculation. The basis for performing IO fusion is that the tensor is split into a suitable size not exceeding the L2 capacity. That is, first fuse some operator nodes into an entire fusion group, and then continuously adjust the tensor splitting size and record the benefit. The benefit of a single operator node is the basis of the search. The splitting benefit of the splitting and fusion scheme is calculated based on the fusion group. The process of the scheme search is to find all possible fusion groups and score them, and find the one with the highest score. Suppose there are n operator nodes in a search area. When using the dynamic programming search method for search, the specific process is as follows:
[0035] The first step: Establish a benefit calculation method for the fusion group from operator node x to operator node y as a basic function. For the fusion group from operator node x to operator node y, first calculate what the storage peak occupied by the original tensor size is and the quantitative relationship with the L2 storage capacity, and then divide it into at least this number. According to this number of divisions, update the tensor sizes within the entire fusion group, recalculate the storage usage peak. If it is higher than the L2 storage capacity, continue to increase the number of divisions, reduce the tensor size until the tensor after division can be cached by the L2, calculate the benefit of the fusion group at this time and record it. According to the preset optimization search method, continue to calculate and record the benefits under other divisions, such as aligning the division to the SIP number for better hardware performance, so as to select the best division method and record it as B(x, y). The second step: Let x = 1 and y = 1, that is, when there is only one operator node, then B(1, 1) will be recorded as the best solution O(1, 1) from operator node 1 to operator node 1. The third step: Let x = 1 and y = 2. Then O(1, 1) is known. Calculate B(1, 2) and B(2, 2) according to the above step one. Compare which is larger between O(1, 1) + B(2, 2) and B(1, 2). The solution with the maximum benefit is the best solution from operator node 1 to operator node 2, denoted as O(1, 2). The fourth step: And so on. According to the state transition equation: O(1, n) = Max{ B(1, n), O(1, 1)+B(2, n), O(1, 2)+B(3, n), O(1, 3)+B(4, n),...., O(1, k)+B(k + 1, n), ……, O(1, n - 3)+B(n - 2, n), O(1, n - 2)+B(n - 1, n), O(1, n - 1)+B(n, n)}, continue to calculate O(1, 3), O(1, 4)... until O(1, n), and then the best division and fusion solution for the current search area with n operator nodes can be obtained. For example, when the search area includes five operator nodes ABCDE, during the dynamic programming search process, it will search for division and fusion methods such as A + B + C + D + E, AB + C + D + E, ABC + D + E, A + BC + D + E, ABCD + E, AB + CD + E, A + BCD + E, A + B + C + DE, AB + C + DE, ABC + DE, A + BC + DE, AB + CDE, A + BCDE, ABCDE, etc. Each fusion group will be scored according to the L2 capacity size to obtain the division and the division benefit at that time, so as to select the best division and fusion solution in the search area. In addition, since the dynamic programming search method can make full use of the already searched solutions without calculating all the search solutions, the search efficiency can be improved.For example, when searching for the optimal segmentation and fusion scheme from A to D, assuming that the benefit of fusing ABCD is greater than the non-fusion benefit of A + B + C + D, then when continuing to search from A to E, the previous search results can be fully utilized. That is, when searching for the two schemes of [ABCD] + E and [A + B + C + D] + E, there is no need to calculate obviously, because the benefit of the first part in the former scheme is the same as that of ABCD, and the benefit of the first part in the latter scheme is the same as that of [A + B + C + D], and the latter is definitely less than the former. Therefore, in this embodiment, search optimization will be performed according to the existing information during the search process, only necessary searches are carried out, but still the global optimal segmentation and fusion scheme can be ensured to be searched for.
[0036] Step S103, when it is determined that the optimal segmentation and fusion scheme includes a fusion group, convert the fusion group into a cyclic semantic node.
[0037] Optionally, converting the fusion group into a cyclic semantic node includes: determining the sub-operator nodes obtained by segmenting each operator node in the fusion group according to the segmentation times; inserting slice nodes and anti-slice nodes representing tensor transfer between the highest-level storage layer and the shared storage layer in the fusion group, where the number of anti-slice nodes is the same as the segmentation times; combining the sub-operator nodes, slice nodes and anti-slice nodes to construct sub-fusion groups, where the sub-fusion groups include sub-operator nodes of different types, and the number of sub-fusion groups is the same as the segmentation times; connecting each sub-fusion group with connection nodes to obtain an expanded fusion group, and converting the expanded fusion group into a cyclic semantic node, where the sub-fusion groups are respectively used to process sub-tensor data of the same size, and the sum of each sub-tensor data is the original processed tensor data of the fusion group.
[0038] Optionally, converting the expanded fusion group into a cyclic semantic node includes: identifying the calculation logics of all slices in the expanded fusion group to obtain slices with the same calculation logic, where the offset information of each slice with the same calculation logic is different; combining the slices with the same calculation logic into a cyclic slice with cyclic semantics, where the number of cycles of the cyclic slice is the same as the segmentation times; converting the expanded fusion group into a cyclic semantic node based on the cyclic slice.
[0039] Specifically, in this embodiment, when it is determined that the optimal segmentation and fusion scheme includes a fusion group, if the operator nodes in the fusion group are processed in a fully expanded manner during tensor data segmentation, it will cause one operator node on the computational graph to become multiple, thereby increasing the complexity of the computational graph, and the compilation processing time of the entire network will become longer. In this embodiment, after converting the fusion group into a loop semantic node, since the number of loops can be marked in the loop semantic node, only the operator nodes involved in the processing flow of the tensor data for a single time need to be expressed in terms of structural expression. The loop semantic node will be automatically scheduled multiple times during operation, thus achieving the same effect as global expansion. This loop semantic node can be called a while node.
[0040] In a specific implementation, as Figure 3 shown in the schematic diagram of the conversion process of the loop semantic node. When it is determined that a fusion group includes three operator nodes A, B, and C, and the highest dimension of the tensor data originally processed by the fusion group is bs6, when the segmentation times for this fusion group are determined to be 3 according to the optimal segmentation and fusion scheme, each operator node can be split into three sub-operator nodes, and the tensor data processed by each sub-operator node is reduced to bs2. For example, operator node A is split into three sub-operator nodes a, operator node B is split into three sub-operator nodes b, and operator node C is split into three sub-operator nodes c, and the functions of the three sub-operator nodes split from each operator node are the same. Of course, in this embodiment, it is only an example, and the number of sub-operator nodes split from each operator node is not limited. In this embodiment, a slice node slice for tensor data transfer from L3 to L2 will be inserted at the beginning of the fusion group, and an inverse slice node deslice for tensor data transfer from L2 to L3 will be added at the end of the fusion group, and the number of slice nodes and inverse slice nodes is 3, that is, the same as the segmentation times of each operator node in the fusion group. Additionally, Figure 3Only the linear structure fusion group with single input and single output is taken as an example for illustration. In actual applications, the fusion group can also be a tree structure with multiple inputs and single output. In this case, the number of anti-slice nodes is still the same as the number of slicing times, but the number of slice nodes is related to the input operator nodes and the number of slicing times. Therefore, the number of slice nodes and anti-slice nodes is no longer the same. Of course, only examples are given in this embodiment, and the specific number of slice nodes is not limited. In this embodiment, a sub-operator node is extracted from the sub-operator nodes split from each operator node and connected in the execution order. Then, a slice node is connected to the first operator node, and an anti-slice node is connected to the last operator node to construct a sub-fusion group. For example, slice-a-b-c-deslice. Therefore, each sub-fusion group includes sub-operator nodes of different types. Since each operator node can be split into three identical sub-operator nodes, three sub-fusion groups can be obtained by using the above construction method, and the obtained sub-fusions will be connected by connection nodes such as concat to obtain an expanded fusion group. From Figure 3 It can be seen that the operator nodes are split due to tensor slicing in the expanded fusion group, and the number of nodes increases significantly. Each sub-fusion group in this embodiment processes bs2 tensor data respectively. When the operator nodes in the same sub-fusion group perform operations, the output of the previous sub-operator node is not written back to L3, but is temporarily stored on L2 as the input for the subsequent operator nodes until the last sub-operator node writes it back to L3. For example, after the a operator node loads data from L3 and processes it to obtain the first processing result, the first processing result is not written back to L3 but is directly transmitted to the b operator node and used as the input for the b operator node. Then, the b operator node processes it according to the first processing result to obtain the second processing result, and the second processing result is also not written back to L3 but is directly transmitted to the c operator node and used as the input for the c operator node. Finally, the c operator node processes it according to the second processing result to obtain the third processing result and finally writes the third processing result back to L3. As can be seen, the operation processes of the sub-operator nodes in the same sub-fusion group are all executed on L2, and only the data transfer operations related to L3 are involved when obtaining the initial data at the beginning and writing back the final processing result. Therefore, compared with the operation method of writing back once for each processing in the prior art, the number of data transmissions is significantly reduced. Since the structure, processing logic, and size of the processed tensor data of each sub-fusion group are the same, the expanded fusion group can be converted into a loop semantic node.
[0041] Among them, the expanded fusion group is converted into a loop semantic node, that is, multiple structures, processing logics, and sub-fusion groups with the same size of processed tensor data in the fusion group are merged into a while loop semantic node. Specifically in the specific operation, the slice nodes with the same calculation logic in the expanded fusion group are converted into loop slices with loop semantics. For example, three slice nodes are converted into one loopslice loop slice, and three deslice reverse slices are converted into one loopdeslice loop reverse slice. And the loop counts of loopslice and loopdeslice are both 3. The corresponding three sub-fusion groups are merged into one, and loopslice and loopdeslice are used for connection to obtain the while loop semantic node. The input and output of the while loop semantic node are complete bs6 tensor data, and the tensor data processed inside each loop is bs2. Thus, the number of operator nodes in the expanded fusion group is reduced through the loop processing method. And in this embodiment, the concat at the L3 level in the expanded fusion group is optimized to avoid actually performing multiple L3 splicing transports. The L3 output of the while loop semantic node only allocates a storage space of the size of the original tensor. After each slice loop is executed, the small-scale calculation result tensor data on L2 each time is directly transported from L2 to the corresponding position of L3 output by using different offsets, thus avoiding the transport of tensors on L3.
[0042] Step S104, modify the computation graph according to the loop semantic nodes obtained in each search area.
[0043] Optionally, modifying the computation graph according to the loop semantic nodes obtained in each search area includes: determining the specified positions of each fusion group on the computation graph; replacing the fusion group with the matching loop semantic nodes at the specified positions.
[0044] Specifically, in this embodiment, after obtaining the fusion groups in all search areas of the computation graph and converting the fusion groups into loop semantic nodes, the specified positions of each fusion group in the computation graph will be determined, and the operator nodes at this position will be replaced with the loop semantic nodes matched by the fusion group, so as to optimize the computation graph. And the structure of the optimized computation graph is more concise, the internal processing efficiency is higher, and the extra IO transports are avoided, thus significantly improving the overall operation performance of the computation graph.
[0045] In this embodiment, the optimal segmentation and fusion scheme is searched by adopting dynamic programming search for the computational graph, so as to maximize the use of historical search results to simplify the computational complexity in the search process. And in the optimal fusion scheme, the operator nodes in the same fusion group only perform operator calculations in the shared storage layer, thus avoiding the transmission of tensor data between the high-level storage layer and the low-level storage layer, improving the data processing efficiency of the computational graph, and further improving the computational performance of the computational graph by converting the complex fusion group into a simple loop semantic node.
[0046] Embodiment 2
[0047] Figure 4 The flowchart of a compilation optimization method for a computational graph provided by Embodiment 2 of the present invention. This embodiment is based on the above embodiment, and specifically illustrates the segmentation benefits of each segmentation and fusion scheme in each search area during the dynamic programming search process in the above embodiment, as Figure 4 shown. The method includes:
[0048] Step S201, determine whether the segmentation scheme includes a fusion group. If so, execute Step S202; otherwise, execute Step S204.
[0049] Specifically, in this embodiment, the segmentation benefits of the segmentation and fusion scheme are calculated based on the fusion group. If the operator nodes in the calculation are not fused into any fusion group, there is no benefit. The tensor data of all operator nodes in each fusion group adopts the same segmentation method, so that the cut sizes match and the output of the previous operator can be directly given to the next operator as input. Therefore, in this embodiment, it is first necessary to determine whether there is a fusion group in the segmentation and fusion scheme before calculating the segmentation benefits of the segmentation and fusion scheme.
[0050] Step S202, calculate the segmentation benefits of each fusion group.
[0051] Optionally, calculating the segmentation benefits of each fusion group includes: obtaining the peak storage usage of all operator nodes in each fusion group in the shared storage layer, and determining the number of segmentation times for the fusion group according to the peak storage usage and the storage capacity of the shared storage layer; obtaining the operator overhead of each operator node in the fusion group in the highest-level storage layer before segmentation, and the operator overhead in the shared storage layer after segmentation according to the number of segmentation times, where the shared storage layer is located below the highest-level storage layer; determining the segmentation benefits of each operator node in the fusion group in the shared storage layer according to the number of segmentation times, the operator overhead in the highest-level storage layer, and the operator overhead in the shared storage layer; and obtaining the segmentation benefits of the fusion group according to the segmentation benefits of each operator node in the shared storage layer.
[0052] Optionally, obtain the splitting benefit of the fusion group according to the splitting benefits of each operator node in the shared storage layer, including: obtaining the transfer overhead between the highest-level storage layer and the shared storage layer of the input operator node and the output operator node in the fusion group; obtaining the sum of the splitting benefits of all operator nodes in the shared storage layer of the fusion group, and taking the difference between the sum of the splitting benefits and the transfer overhead as the splitting benefit of the fusion group.
[0053] In a specific implementation, when there are five operator nodes ABCDE in the computation graph and the original processed tensor data is bs6, when a splitting fusion scheme is obtained during the search: AB is fused, and the tensor is split into bs2; C remains as a single operator in L3 without splitting; DE is fused, and the tensor is split into bs3. The benefit of a single operator only considers the normal overhead when the tensor of the operator itself is in L3, and the total overhead when the tensor is in L2 when split into multiple processes. The benefits of a single operator for various splits, which dimensions to split, how many parts to split into, and the size of each part are the basic information for dynamic programming search, and not being fused means that the operator nodes in the fusion group have no splitting benefits. Therefore, in this embodiment, only the splitting benefits of each operator node in the fusion group are considered, and the splitting benefit of the entire fusion group is obtained based on the splitting benefits of each operator node in the fusion group. Among them, when calculating the splitting benefit of the fusion group AB, the splitting benefits of operator nodes A and B can be obtained respectively. In this embodiment, taking the obtaining of the splitting benefit of operator node A as an example, when obtaining the splitting benefit of operator node A, the peak storage usage of all operator nodes in the fusion group AB in L2 will be obtained first, and the splitting times for the fusion group AB will be determined according to the peak storage usage and the storage capacity of L2. For example, the splitting is first performed according to the highest dimension. If it still exceeds the storage size of L2 after splitting the highest dimension, then continue to split the second-highest dimension to ensure continuity. Then, the following formula (1) is used to calculate the splitting benefit of operator node A when executed on L2 after splitting:
[0054] Splitting benefit of a single operator node = L3 operator overhead - number of splitting times * L2 operator overhead (1)
[0055] Among them, the L3 operator overhead is the overhead when the tensor data of a single operator node is stored in L3 without splitting, and the L2 operator overhead is the overhead for a single split when the single operator node exists in L2. Since the actual execution strategy of operator nodes at different storage levels and different tensor shapes is generated according to known specific algorithms, including the data size transferred each time between high and low-level storages, the tensor shape processed by the computing unit each time, etc. And the frequency of the AI hardware platform, the data processing capabilities of the computing unit under different operators, the data transfer rate, the task issuance overhead, etc. are also all known. Therefore, combining the above information, the following formula (2) can be used to obtain the overhead of a single operator node at different storage levels:
[0056] Operator overhead = fixed overhead of calling the operator + size of the first input tensor / data transfer rate + max (size of a single input tensor / data processing capacity of the computing unit, size of a single input tensor / data transfer rate + size of a single output tensor / data transfer rate) + size of the tensor of the final result / data transfer rate (2)
[0057] Among them, in this embodiment, after obtaining the splitting benefits of each single-operator node in the fusion group AB, the splitting benefit of the fusion group AB can be obtained based on the splitting benefits of each operator node in the shared storage layer. When calculating the splitting benefit of the fusion group, since the fusion group composed of multiple operator nodes only needs to be called once during execution, eliminating the fixed operator call overhead multiple times and obtaining the basic benefit. Therefore, when calculating the benefit of a certain splitting of the fusion group, in addition to accumulating the splitting benefits of all internal operators, it is necessary to deduct the transfer overhead of the last output operator node for each tensor split from L2 to L3, and the transfer overhead of each input operator node from L3 to L2 after each split, and combine the transfer overheads in the above two directions into the transfer overhead between L3 and L2. In addition, the sum of the splitting benefits of all operator nodes in the fusion group at L2 will be obtained, and the splitting benefit of the fusion group will be obtained based on the following formula (3):
[0058] Fusion group splitting benefit = sum of splitting benefits of all operator nodes - transfer overhead between L3 and L2 (3)
[0059] It is worth mentioning that in this embodiment, when the operator node is split into multiple executions, the execution strategy of the operator will change, and due to the hardware characteristics affecting the transfer efficiency of tensors of different shapes, the overhead of computing unit calculation and tensor transfer may become more compared to not doing the split, and there may be factors such as multiple repeated transfers of weights. Splitting a single operator node does not necessarily result in a positive benefit. However, when multiple consecutive operator nodes are fused, even if the splitting benefit of a certain single-operator node in the middle is negative, it may still maximize the benefit of the entire fusion group. For example, if the splitting benefits of 5 consecutive operator nodes are 2, 2, 2, -1, 2, then obviously fusing the fourth operator node with a negative benefit together with the fifth operator node will obtain more benefits, rather than only fusing the first three operator nodes. The actual situation is more complex, and in this embodiment, the global optimal splitting scheme can be searched through the introduced dynamic programming search technology to obtain the best splitting and fusion scheme.
[0060] Step S203, obtain the total splitting benefits of all fusion groups in each splitting scheme, and use the total splitting benefits as the splitting benefits of each splitting and fusion scheme.
[0061] Specifically, for the splitting and fusion scheme: AB fusion, the tensor is split to bs2; C remains as a single operator of L3 without splitting; for DE fusion, when the tensor is split to bs3, based on the sum of the splitting benefits of the fusion group AB and the splitting benefits of the fusion group DE obtained in the above manner, the splitting benefits of the fusion group AB and the fusion group DE can be added to obtain the total splitting benefit, and the total splitting benefit is used as the splitting benefit of this splitting scheme. Of course, in this embodiment, only the process of obtaining the splitting benefit of one splitting and fusion scheme is taken as an example for illustration. The process of obtaining the splitting benefits of other splitting and fusion schemes is roughly the same, and will not be elaborated in this embodiment.
[0062] Step S204, directly determine that the splitting benefit of the splitting and fusion scheme is 0.
[0063] In this embodiment, the optimal splitting and fusion scheme is searched for the computational graph by using the dynamic programming search method, so as to maximize the use of historical search results to simplify the computational amount in the search process. And in the optimal fusion scheme, the operator nodes in the same fusion group only perform operator calculations in the shared storage layer, thus avoiding the transmission of tensor data between the high-level storage layer and the low-level storage layer, improving the data processing efficiency of the computational graph, and further improving the computational performance of the computational graph by converting the fusion group with complex structure into simple loop semantic nodes.
[0064] Embodiment III
[0065] Figure 5 It is a schematic structural diagram of a compilation optimization device for a computational graph provided by Embodiment III of the present invention. As Figure 5 shown, the device includes: a search area division module 310, an optimal splitting and fusion scheme search module 320, a loop semantic node conversion module 330, and a computational graph modification module 340.
[0066] Among them, the search area division module 310 is used to divide the computational graph running on the multi-level storage chip to obtain multiple search areas;
[0067] The optimal splitting and fusion scheme search module 320 is used to search for the optimal splitting and fusion scheme for each search area by using the dynamic programming search method;
[0068] The loop semantic node conversion module 330 is used to convert the fusion group into a loop semantic node when it is determined that the optimal splitting and fusion scheme includes a fusion group, where all operator nodes in the fusion group perform tensor transfer in the shared storage layer of the multi-level storage chip;
[0069] The computational graph modification module 340 is used to modify the computational graph according to the loop semantic nodes obtained in each search area.
[0070] Optionally, a search area division module is configured to identify specified operator nodes that do not support splitting from each operator node in the computation graph, where the specified operator nodes include global normalization operator nodes and user-defined operator nodes;
[0071] Filter the specified operator nodes from the computation graph to obtain a processed computation graph;
[0072] Perform area division on the processed computation graph to obtain multiple search areas, where each search area includes multiple connected operator nodes.
[0073] Optionally, the optimal splitting and fusion scheme search module includes:
[0074] A splitting and fusion scheme search unit is configured to search for all splitting and fusion schemes for operator nodes in each search area by using a dynamic programming search method;
[0075] An optimal splitting and fusion scheme determination unit is configured to calculate the splitting benefit of each splitting and fusion scheme in each search area, and use the splitting and fusion scheme with the maximum splitting benefit as the optimal splitting and fusion scheme in each search area.
[0076] Optionally, the optimal splitting and fusion scheme determination unit includes a splitting benefit calculation unit for the splitting and fusion scheme, which is configured to directly determine that the splitting benefit of the splitting and fusion scheme is 0 when it is determined that the splitting and fusion scheme does not include a fusion group;
[0077] When it is determined that the splitting and fusion scheme includes a fusion group, calculate the splitting benefit of each fusion group, obtain the total splitting benefit of all fusion groups in each splitting scheme, and use the total splitting benefit as the splitting benefit of each splitting and fusion scheme.
[0078] Optionally, the splitting benefit calculation unit for the splitting and fusion scheme is configured to obtain the peak storage usage of all operator nodes in each fusion group in the shared storage layer, and determine the number of splitting times for the fusion group according to the peak storage usage and the storage capacity of the shared storage layer;
[0079] Obtain the operator overhead of each operator node in the fusion group before splitting when the tensor is stored in the highest-level storage layer, and the operator overhead in the shared storage layer after splitting according to the number of splitting times, where the shared storage layer is located below the highest-level storage layer;
[0080] Determine the splitting benefit of each operator node in the fusion group in the shared storage layer according to the number of splitting times, the operator overhead of the highest-level storage layer, and the operator overhead of the shared storage layer;
[0081] Obtain the splitting benefit of the fusion group according to the splitting benefit of each operator node in the shared storage layer.
[0082] Optionally, the segmentation benefit calculation unit of the segmentation and fusion scheme is further configured to obtain the transfer overhead between the input operator node and the output operator node in the fusion group between the highest-level storage layer and the shared storage layer;
[0083] Obtain the sum of the segmentation benefits of all operator nodes in the fusion group in the shared storage layer, and use the difference between the sum of the segmentation benefits and the transfer overhead as the segmentation benefit of the fusion group.
[0084] Optionally, the loop semantic node conversion module is configured to determine the sub-operator nodes obtained by segmenting each operator node in the fusion group according to the number of segmentation times;
[0085] Insert slice nodes and anti-slice nodes representing tensor transfer between the highest-level storage layer and the shared storage layer in the fusion group, where the number of anti-slice nodes is the same as the number of segmentation times;
[0086] Combine the sub-operator nodes, slice nodes, and anti-slice nodes to construct sub-fusion groups, where the sub-fusion groups include sub-operator nodes of different types, and the number of sub-fusion groups is the same as the number of segmentation times;
[0087] Connect each sub-fusion group using connection nodes to obtain an expanded fusion group, and convert the expanded fusion group into a loop semantic node, where the sub-fusion groups are respectively used to process sub-tensor data of the same size, and the sum of all sub-tensor data is the original processed tensor data of the fusion group.
[0088] Optionally, the loop semantic node conversion module is further configured to identify the calculation logic of all slices in the expanded fusion group to obtain slices with the same calculation logic, where the offset information of each slice with the same calculation logic is different;
[0089] Merge the slices with the same calculation logic into a loop slice with loop semantics, where the number of loops of the loop slice is the same as the number of segmentation times;
[0090] Convert the expanded fusion group into a loop semantic node based on the loop slice.
[0091] Optionally, the computation graph modification module is configured to determine the specified positions of each fusion group on the computation graph; replace the fusion group with a matching loop semantic node at the specified positions.
[0092] The computation graph compilation optimization device provided by the embodiments of the present invention can execute the computation graph compilation optimization method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0093] Embodiment 4
[0094] Figure 6 It is a schematic structural diagram of a computer device provided for Embodiment 4 of the present invention, asFigure 6 As shown, the computer device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the computer device can be one or more, Figure 6 and one processor 610 is taken as an example herein; the processor 610, the memory 620, the input device 630, and the output device 640 in the computer device can be connected through a bus or other means, Figure 6 and connection through a bus is taken as an example herein.
[0095] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the compilation optimization method of the computational graph in the embodiments of the present invention. The processor 610 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 620, that is, implements the above-mentioned compilation optimization method of the computational graph.
[0096] Partition the computational graph running on the multi-level storage chip to obtain multiple search regions;
[0097] For each search region, adopt a dynamic programming search method to search for the optimal segmentation and fusion scheme;
[0098] When it is determined that the optimal segmentation and fusion scheme includes a fusion group, convert the fusion group into a loop semantic node, where each operator node in the fusion group performs operator operations in the shared storage layer of the multi-level storage chip;
[0099] Modify the computational graph according to the loop semantic nodes obtained in each search region.
[0100] The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 620 may include a high-speed random access memory, and may further include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely set relative to the processor 610, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0101] The input device 630 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device. The output device 640 may include a display device such as a display screen.
[0102] Embodiment Five
[0103] Embodiment 5 of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a compilation optimization method for a computational graph when executed by a computer processor;
[0104] Partition the computational graph running on a multi-level storage chip to obtain multiple search regions;
[0105] For each search region, adopt a dynamic programming search method to search for the best splitting and fusion scheme;
[0106] When it is determined that the best splitting and fusion scheme includes a fusion group, convert the fusion group into a loop semantic node, where each operator node in the fusion group performs operator operations in the shared storage layer of the multi-level storage chip;
[0107] Modify the computational graph according to the loop semantic nodes obtained in each search region.
[0108] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the above method operations, and can also execute related operations in the compilation optimization method of the computational graph provided by any embodiment of the present invention.
[0109] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0110] It should be noted that in the embodiments of the above-mentioned compilation optimization device for a computational graph, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0111] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for compiling and optimizing a computational graph, characterized in that: include: The computation graph running on the multi-level storage chip is divided to obtain multiple search areas; A dynamic programming search method is used to search for each of the search areas to obtain the best segmentation and fusion solution; When it is determined that the optimal splitting and fusion scheme includes a fusion group, converting the fusion group into a loop semantic node, wherein all operator nodes in the fusion group perform tensor transport in a shared storage layer of the multi-level storage chip; Modifying the computation graph according to the loop semantic nodes obtained in each of the search areas; The converting the fusion group into a cyclic semantic node includes: determining sub-operator nodes obtained by splitting each operator node in the fusion group according to a number of splits, wherein the number of splits is determined according to a storage usage peak of all operator nodes in the fusion group in the shared storage layer and a storage capacity of the shared storage layer; Inserting a slice node and an anti-slice node representing tensor transfer between the highest-level storage layer and the shared storage layer into the fusion group, wherein the number of the anti-slice nodes is the same as the number of slicing times; Combining the sub-operator nodes, the slice nodes and the anti-slice nodes to construct a sub-fusion group, wherein the sub-fusion group includes the sub-operator nodes of different types, and the number of the sub-fusion groups is the same as the number of segmentation times; Each of the sub-fusion groups is connected by a connection node to obtain an expanded fusion group, and the expanded fusion group is converted into the loop semantic node, wherein the sub-fusion groups are respectively used to process sub-tensor data of the same size, and the sum of each of the sub-tensor data is the original processed tensor data of the fusion group.
2. The method according to claim 1, characterized in that: The step of dividing the computation graph running on the multi-level storage chip to obtain multiple search areas includes: Identify each operator node in the computation graph to obtain a designated operator node that does not support segmentation, wherein the designated operator node includes a global normalization operator node and a user-defined operator node; Filtering the designated operator node from the computation graph to obtain a processed computation graph; The processed computation graph is divided into regions to obtain a plurality of search regions, wherein each of the search regions includes a plurality of connected operator nodes.
3. The method according to claim 1, characterized in that: The method of using a dynamic programming search method to search for each search area to obtain the best segmentation and fusion solution includes: A dynamic programming search method is used to search for each of the search areas to obtain all the segmentation and fusion solutions for the operator nodes; The segmentation benefit of each segmentation and fusion scheme in each search area is calculated, and the segmentation and fusion scheme with the largest segmentation benefit is used as the optimal segmentation and fusion scheme in each search area.
4. The method according to claim 3, characterized in that The calculating the segmentation benefit of each segmentation and fusion scheme in each search area includes: When it is determined that the splitting and fusion scheme does not include a fusion group, the splitting benefit of the splitting and fusion scheme is directly determined to be 0; When it is determined that the splitting and fusion scheme includes a fusion group, the splitting benefit of each fusion group is calculated, the total splitting benefit of all fusion groups in each splitting and fusion scheme is obtained, and the total splitting benefit is used as the splitting benefit of each splitting and fusion scheme.
5. The method according to claim 4, characterized in that The calculating the splitting benefits of each fusion group includes: Obtaining the storage usage peak of all operator nodes in each of the fusion groups in the shared storage layer, and determining the number of splits for the fusion group according to the storage usage peak and the storage capacity of the shared storage layer; Obtaining the operator overhead of the tensor stored in the highest-level storage layer before each operator node in the fusion group is split, and the operator overhead in the shared storage layer after being split according to the number of splits, wherein the shared storage layer is located at the lower layer of the highest-level storage layer; Determine the slicing benefit of each operator node in the fusion group in the shared storage layer according to the number of slicing times, the operator overhead of the highest-level storage layer, and the operator overhead of the shared storage layer; The splitting benefit of the fusion group is obtained according to the splitting benefit of each operator node in the shared storage layer.
6. The method according to claim 5, characterized in that The obtaining the slicing benefit of the fusion group according to the slicing benefit of each operator node in the shared storage layer includes: Obtaining the transport overhead of the input operator nodes and the output operator nodes in the fusion group between the highest-level storage layer and the shared storage layer; Obtain the sum of the split benefits of all operator nodes in the fusion group in the shared storage layer, and convert the The difference between the splitting benefit and the transport overhead is used as the splitting benefit of the fusion group.
7. The method according to claim 1, characterized in that The converting the expanded fusion group into the loop semantic node comprises: Performing calculation logic identification on all slices in the expanded fusion group to obtain slices with the same calculation logic, wherein the offset information of the slices with the same calculation logic is different; Merge the slices with the same computational logic into a loop slice with loop semantics, wherein the number of loops of the loop slice is the same as the number of slices; The expanded fusion group is converted into the loop semantic node based on the loop slice.
8. The method according to claim 1, characterized in that The modifying the computation graph according to the loop semantic nodes obtained in each of the search areas includes: Determine a designated position of each of the fusion groups on the computational graph; The fusion group is replaced at the specified position with the matching loop semantic node.
9. A computation graph compilation optimization device, characterized in that: include: A search area partitioning module is used to partition the computational graph running on the multi-level storage chip to obtain multiple search areas; The optimal segmentation and fusion solution search module is used to search for the optimal segmentation and fusion solution by using a dynamic programming search method for each search area; A loop semantic node conversion module, used for converting the fusion group into a loop semantic node when it is determined that the optimal segmentation and fusion scheme includes a fusion group, wherein all operator nodes in the fusion group perform tensor transport in the shared storage layer of the multi-level storage chip; A computation graph modification module, used to modify the computation graph according to the loop semantic nodes obtained in each of the search areas; The cyclic semantic node conversion module is used to determine the child operator nodes obtained by splitting each operator node in the fusion group according to the number of splits, wherein the number of splits is determined according to the storage usage peak of all operator nodes in the fusion group in the shared storage layer and the storage capacity of the shared storage layer; Inserting a slice node and an anti-slice node representing tensor transfer between the highest-level storage layer and the shared storage layer into the fusion group, wherein the number of the anti-slice nodes is the same as the number of slicing times; Combining the sub-operator nodes, the slice nodes and the anti-slice nodes to construct a sub-fusion group, wherein the sub-fusion group includes the sub-operator nodes of different types, and the number of the sub-fusion groups is the same as the number of segmentation times; Each of the sub-fusion groups is connected by a connection node to obtain an expanded fusion group, and the expanded fusion group is converted into the loop semantic node, wherein the sub-fusion groups are respectively used to process sub-tensor data of the same size, and the sum of each of the sub-tensor data is the original processed tensor data of the fusion group.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
11. A computer executable instruction storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Calculation graph processing method and device, readable medium and electronic equipment
CN114841327A
Automatic operator fusion method of computational graph and related product
CN115756478A
Deep learning compilation optimization method and device, electronic equipment and storage medium
CN118780351A