A Compilation Optimization Method and Device for a Computation Graph Based on Multi-Level Storage Chips
By filtering and dynamic type acquisition of candidate operator nodes in the calculation graph, fused operator nodes and tensor handling are obtained in the shared storage layer, the problem of low processing efficiency of computing graph data on multi-level memory chips is solved, and the computing performance is improved.
Patent Information
- Application Number
- CN202510300602.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-14
AI Technical Summary
On a multi-level memory chip, the data access efficiency of the calculation graph is low due to frequent writing and reading of intermediate result data during data processing, which significantly reduces the data processing efficiency and calculation performance of the calculation graph.
By filtering and dynamic type acquisition of candidate operator nodes in the calculation graph, forward search and backward search are performed to obtain fusion operator nodes, and non-constant input tensors and output tensors are transported in the shared storage layer of the multi-level memory chip to avoid data transfer between different storage layers.
It effectively reduces the number of tensor data handling times between different storage layers, and improves the data processing efficiency and computing performance of the calculation graph.
Smart Images

Figure CN119806543B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to computer technology, and in particular, to a compilation optimization method and device for a computation graph based on a multi-level storage chip. Background Art
[0002] On a multi-level storage chip, the farther the memory is from the computing unit in terms of hierarchy, the larger its memory capacity and the lower the data access efficiency. A computation graph is a data structure representing the computing process of a neural network. The computing process consists of one operator after another, and there are parts with data dependencies between operators, and data transfer will occur on the outermost layer of storage of the multi-level storage chip.
[0003] When the computation graph processes data, during the computing process of a single operator, the data is read from the outermost layer of storage and transferred to the storage closest to the computing unit, then loaded into the register for computing. After the computing is completed, it is written back to the innermost layer of storage and finally transferred to the outermost layer of storage. On the inference side, the dependent data between operators only needs to reside in the storage temporarily, and the final computing result does not use this part of the data, that is, the intermediate result data. When the intermediate result data is transferred through the outer-level storage, its write-back and read-out access rates are low, thus significantly reducing the data processing efficiency of the computation graph and affecting the computing performance of the computation graph. Summary of the Invention
[0004] The embodiments of the present invention provide a compilation optimization method for a computation graph based on a multi-level storage chip to improve the data processing efficiency and overall computing performance of the computation graph.
[0005] In a first aspect, the embodiments of the present invention provide a compilation optimization method for a computation graph, including: screening a computation graph running on a multi-level storage chip to obtain candidate operator nodes;
[0006] obtaining the dynamic type of each current candidate operator node, where the dynamic type includes single-input single-output, single-input multi-output, multi-input single-output, and multi-input multi-output;
[0007] performing forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic type to obtain fused operator nodes, and updating the computation graph according to the fused operator nodes to obtain an optimized computation graph,
[0008] where all candidate operator nodes in the fused operator nodes perform the transfer of non-constant input tensors and output tensors on the shared storage layer of the multi-level storage chip.
[0009] In a second aspect, the embodiments of the present invention provide a compilation optimization device for a computation graph, including: a candidate operator node screening module for screening a computation graph running on a multi-level storage chip to obtain candidate operator nodes;
[0010] A dynamic type acquisition module for acquiring the dynamic types of current candidate operator nodes, where the dynamic types include single-input single-output, single-input multi-output, multi-input single-output, and multi-input multi-output;
[0011] An optimized computation graph acquisition module for performing forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic types to obtain fused operator nodes, and updating the computation graph according to the fused operator nodes to obtain an optimized computation graph,
[0012] where all candidate operator nodes in the fused operator nodes transfer non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip.
[0013] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method described above when executing the program.
[0014] In a fourth aspect, an embodiment of the present invention provides a storage medium storing computer-executable instructions, on which a computer program is stored, wherein the program implements the method described above when executed by a processor.
[0015] The present invention screens out candidate operator nodes that meet the fusion conditions and the dynamic types of each candidate operator node from the computation graph, and searches for fused operator nodes according to the dynamic types. Since the candidate operator nodes located in the same fused operator node only transfer tensors in the shared storage layer, the number of times of tensor data transfer between different storage layers is avoided, and the data processing efficiency and computing performance of the computation graph are improved. Description of the Drawings
[0016] Figure 1 is a flowchart of a computation graph compilation optimization method provided in Embodiment 1 of the present invention;
[0017] Figure 2 is a schematic structural diagram of a computation graph provided in Embodiment 1 of the present invention;
[0018] Figure 3 is a flowchart of a computation graph compilation optimization method provided in Embodiment 2 of the present invention;
[0019] Figure 4 is a schematic structural diagram of a computation graph compilation optimization device provided in Embodiment 3 of the present invention;
[0020] Figure 5 is a schematic structural diagram of a computer device provided in Embodiment 4 of the present invention. Detailed Embodiments
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention and not for limiting the present invention. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the present invention rather than all the structures are shown in the accompanying drawings.
[0022] Embodiment 1
[0023] Figure 1 As shown in the flowchart of a compilation optimization method for a computational graph provided in Embodiment 1 of the present invention, this embodiment is applicable to the situation of compiling and optimizing a computational graph. This method can be executed by a compilation optimization device for a computational graph, and this device can be implemented in the form of hardware and / or software. Figure 1 As shown, the method includes:
[0024] Step S101, screening the computational graph running on a multi-level storage chip to obtain candidate operator nodes.
[0025] Specifically, the computational graph of this embodiment may include multiple operator nodes. The operator nodes represent operation logic calculation operations such as add and mul, or data transfer operations such as slice. And the computational graph in this embodiment runs on a multi-level storage chip. For example, on a multi-level storage chip including three storage layers, L1, L2, and L3. Among them, the L3 storage is the global storage space, which is the highest-level storage layer, with the largest capacity but the slowest speed; the L2 storage is the shared storage layer, which can be accessed by multiple computing units, such as SIP, with a smaller capacity but a higher speed; the L1 storage is the exclusive storage of the computing unit, with the smallest capacity but the fastest speed. And in this embodiment, the fusion of operator nodes is mainly carried out in the shared storage layer. Among them, the fusion mainly refers to IO fusion. When IO fusion is not performed, the output data of the operator nodes in the computational graph are all default stored in the L3 layer, that is, the highest-level storage layer. When the operator calculates, it loads data from L3 to L2, and then each computing unit SIP loads a part of the data to L1 for calculation. After the calculation is completed, it is written back to L3 from L1 through L2. This operator node is considered to have been executed, and then the next operator node starts to load data from L3. Obviously, this method will increase the number of data transmissions between different storage levels. When IO fusion is adopted, the output of the previous operator node will not be written back to L3, but temporarily stored in L2. Then the next operator node can directly load data from L2, that is, it is only shown that the previous operator node loads data from L3, and the next operator node writes the calculation result after continuously executing the logic of these two operators back to L3. At this time, these two operator nodes are considered to be fused into a fused operator node. Of course, in this embodiment, only a multi-level storage chip including three storage layers is taken as an example for illustration, and the specific form of the multi-level storage chip is not limited.
[0026] Optionally, screening the computational graph running on the multi-level storage chip to obtain candidate operator nodes includes: obtaining the residence status of each operator node in the computational graph in the shared storage layer, where the residence status includes meeting the residence condition and not meeting the residence condition; querying the support status of the input or output of each operator node in the computational graph in the shared storage layer by calling the query interface, where the support status includes support and non-support; and taking the operator nodes with the residence status of meeting the residence condition and the support status of support as candidate operator nodes.
[0027] Optionally, obtaining the residence status of each operator node in the computational graph in the shared storage layer includes: obtaining the tensor size of each operator node in the computational graph and the storage capacity of the shared storage layer; determining whether the tensor size of each operator node is less than the storage capacity. If so, determining that the residence status of the operator node in the shared storage layer meets the residence condition, otherwise, determining that the residence status of the operator node in the shared storage layer does not meet the residence condition.
[0028] Specifically, in this embodiment, before fusion, candidate operator nodes in the computational graph that meet the fusion conditions need to be screened out. When screening, it is mainly to determine whether the operator node meets the conditions for residing in the shared storage layer and whether it supports input or output in the shared storage layer. If both of the above two conditions are met, the operator node is determined to be a candidate operator node. Among them, when judging whether the operator node meets the conditions for residing in the shared storage layer, the tensor size of the operator node and the storage capacity of the shared storage layer need to be obtained. When the tensor size is smaller than the storage capacity, it means that the storage capacity of the shared storage layer can carry the calculation process of the operator node. At this time, it means that the residing state of the operator node in the shared storage layer meets the residing conditions; when the tensor size is larger than the storage capacity, it means that the storage capacity of the shared storage layer is insufficient to carry the calculation process of the operator node. At this time, it means that the residing state of the operator node in the shared storage layer does not meet the residing conditions. Since this application performs fusion in the shared storage layer, it is necessary to screen out the operator nodes that meet the residing conditions in the shared storage layer. In addition, when judging whether the operator node supports input or output in the shared memory layer, it is mainly queried by calling the query interface. Since the support status of the input or output of each operator node in the shared storage layer has been pre-marked, it only needs to be queried through the query interface without additional calculation. The support status includes support and non-support. In this embodiment, the operator nodes that meet both of the above two conditions are screened out and marked as candidate operator nodes in the computational graph. The number of candidate operator nodes can be the same as the number of operator nodes included in the computational graph, that is, all operator nodes in the computational graph meet the fusion conditions, or the number of candidate operator nodes is less than the number of operator nodes included in the computational graph, that is, some operator nodes in the computational graph meet the fusion conditions. Of course, this embodiment is only an example and does not limit the number of the selected candidate operator nodes.
[0029] It should be noted that in this embodiment, different calculation methods can be adopted according to the requirements of the stored content when calculating the tensor size of each operator node. For example, the sum of the sizes of non-constant input tensors and output tensors can be used as the tensor size of the operator node; or, the minimum of the sizes of non-constant input tensors and output tensors can be used as the tensor size of the operator node. Of course, this embodiment is only an example and does not limit the specific calculation method of the tensor size of each operator node, which can be determined according to different requirements of the stored content.
[0030] Step S102, obtain the dynamic types of the current candidate operator nodes.
[0031] Optionally, obtain the dynamic types of current candidate operator nodes, including: obtaining the number of non-constant input tensors of each current candidate operator node, and using the number of non-constant input tensors as the in-degree of the candidate operator node; obtaining the total number of output tensors of each current candidate operator node, and using the total number of output tensors as the out-degree of the candidate operator node; determining the dynamic types of current candidate operator nodes according to the in-degree and out-degree.
[0032] Specifically, in this embodiment, after determining the candidate operators, the dynamic types of each candidate operator node will also be obtained. Among them, the dynamic types include single-input single-output, single-input multi-output, multi-input single-output, and multi-input multi-output, and the dynamic types of each candidate operator node are determined according to the in-degree and out-degree. Among them, the in-degree is mainly determined according to the input parameters of the candidate operator node, and the input parameters mainly refer to the number of non-constant input tensors; the out-degree is determined according to the output parameters of the candidate operator node, and the output parameters mainly refer to the total number of output tensors, and the out-degree in this embodiment is specifically the number of other operators that depend on the output result of the current candidate operator node. For example, when it is determined that the in-degree of the candidate operator node is 1 and the out-degree is 2, then the dynamic type of the candidate operator node is determined to be single-input multi-output. When the in-degree of the candidate operator node is 2 and the out-degree is 2, then the dynamic type of the candidate operator node is determined to be multi-input multi-output. Of course, only examples are given in this embodiment, and the specific values of the in-degree and out-degree of the candidate operator node are not limited. And in this embodiment, the dynamic types of each candidate operator node can be updated according to the fusion result during the subsequent fusion process.
[0033] It should be noted that only the number of non-constant input tensors is considered when calculating the in-degree in this embodiment, because for operator nodes, the input of constants is generated before the running period. For example, when calculating the addition of two numbers, x + y is addition, and x + 1 is also addition, but when adding 1, 1 is already a determined constant, while y is unknown and the result of y needs to be given by the previous operator node. From this, it can be known that the tensor data of constants does not depend on the results calculated by the previous operator nodes, so there is no dependency edge in the computational graph and it is not included in the in-degree.
[0034] Step S103: Perform forward search and backward search on the candidate operator nodes in the computational graph to obtain fused operator nodes, and update the computational graph according to the fused operator nodes to obtain an optimized computational graph.
[0035] Optionally, forward search and backward search are performed on the candidate operator nodes in the computation graph according to the dynamic type to obtain fused operator nodes, and the computation graph is updated according to the fused operator nodes to obtain an optimized computation graph, including: using a specified operator node in the computation graph as the forward search anchor point, and using a greedy algorithm to search forward and fuse with reference to the dynamic type to obtain at least one fused operator node, where the specified operator node includes a candidate operator node or a fused operator node; updating the computation graph according to the fused operator node to obtain an updated computation graph; using the specified operator nodes with single input and multiple outputs and multiple inputs and multiple outputs in the updated computation graph as the backward search anchor points, and using a greedy algorithm to search backward with reference to the dynamic type to obtain supplementary fused operator nodes; and updating the updated computation graph according to the supplementary fused operator nodes to obtain an optimized computation graph.
[0036] Specifically, after marking the candidate operator nodes on the computation graph in this embodiment, the candidate operator nodes can be traversed to search forward and fuse to obtain fused operator nodes, and the computation graph is updated according to the fused operator nodes to obtain an updated computation graph. After obtaining the updated computation graph, the candidate operator nodes or the fused operator nodes obtained by fusion will be traversed again to search backward to supplement the candidate operator nodes that are missed in the fusion during the forward search process to obtain supplementary fused operator nodes, and the updated computation graph is updated again according to the supplementary fused operator nodes to obtain an optimized computation graph. Through the forward search and the backward supplementary search process, as many candidate operator nodes that meet the fusion requirements as possible can be fused. Among them, all candidate operator nodes in the fused operator nodes perform the transfer of non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip. Therefore, the number of transfers of tensor data between different storage levels in the finally obtained optimized computation graph will be significantly reduced.
[0037] Optionally, using a greedy algorithm to search forward and fuse with reference to the dynamic type to obtain at least one fused operator node includes: searching forward for a single-input single-output or multi-input single-output specified operator node to be fused that has a dependency relationship with the current specified operator node; obtaining the peak memory usage of the specified operator node to be fused and the forward search anchor point in the shared storage layer; and determining the fused operator node using a greedy algorithm based on the peak memory usage and the storage capacity of the shared storage layer.
[0038] Optionally, a greedy algorithm is used to determine the fused operator nodes based on the peak memory usage and the storage capacity of the shared storage layer, including: determining whether the peak memory usage exceeds the storage capacity of the shared storage layer; if so, sequentially adjusting the specified operator nodes to be fused that have no dependencies, and obtaining the fused operator nodes according to the adjustment result; otherwise, fusing the specified operator nodes to be fused with the forward search anchor point to obtain intermediate fused operator nodes, updating the dynamic type of the intermediate fused operator nodes, using the intermediate fused operator nodes as the new forward search anchor point to increase the search range and re-perform the search, and obtaining the fused operator nodes according to the result of the re-search.
[0039] Optionally, obtaining the fused operator nodes according to the adjustment result includes: determining whether there is an adjustment result where the peak memory usage does not exceed the storage capacity of the shared storage layer; if so, obtaining the operator node sequence corresponding to the adjustment result, and fusing the specified operator nodes to be fused and the anchor point according to the operator node sequence to obtain the fused operator nodes; otherwise, abandoning the fusion of the anchor point and the specified operator nodes to be fused.
[0040] In a specific implementation, such as Figure 2The figure shows a schematic structural diagram of a computational graph. The computational graph includes five operator nodes A, B, C, D, and E, and each operator node is determined to meet the fusion conditions through screening. Therefore, the above five operator nodes are all marked as candidate operator nodes. First, when performing forward search, a specified operator node in the computational graph is used as the forward search anchor point, and the specified operator node includes a candidate operator node or a fused operator node. For example, taking A as the anchor point, search forward for a specified operator node to be fused with a single-input single-output or multi-input single-output that has a dependency relationship with A, such as B. Then calculate the peak memory usage a of A and B in the shared storage layer. When it is determined that the peak memory usage a does not exceed the storage capacity of the shared storage layer, directly fuse A and B to obtain an intermediate fused operator node A2(AB). Then expand the search scope again and continue to search for a specified operator node to be fused with a single-input single-output or multi-input multi-output that has a dependency relationship with A2(AB), such as C. Then calculate the peak memory usage b of ABC in the shared storage layer. When it is determined that b exceeds the memory capacity of the shared storage layer, adjust the order of the specified operator nodes to be fused that do not have a dependency relationship. For example, when it is determined that both B and C depend on A, keep the position of A unchanged and adjust the order of C and B. Then calculate the peak memory usage b' of ACB in the shared storage layer. When it is determined that b' does not exceed the memory capacity of the shared storage layer, fuse A2(AB) with C to obtain an intermediate fused operator node A3(ACB). Then expand the search scope again and continue the search to find a specified operator node to be fused with a single-input single-output or multi-input multi-output that has a dependency relationship with A3(ACB); however, when it is determined that b' still exceeds the memory capacity of the shared storage layer, abandon the fusion of A2(AB) and C, take the intermediate fused operator node A2(AB) as the fused operator node obtained by the current search, and take C as a new forward search anchor point to perform the fusion search again in the above manner. For example, after searching the computational graph according to the above forward search method, if the fused operator node obtained includes A2(AB), the updated computational graph obtained by updating the computational graph according to the fused operator node includes four operator nodes A2(AB), C, D, and E.
[0041] It should be noted that within the fused operator node, when executing different operator nodes, the usage amount of the shared storage layer is different, and the peak memory usage is determined by the maximum usage moment. Moreover, the order of the operator nodes determines the temporarily stored data, thereby affecting the maximum moment of the memory usage amount. The peak memory usage corresponding to different operator node orders can be different. Therefore, when the peak memory usage of the fused operator node exceeds the storage capacity of the shared storage layer, the order of the operator nodes that do not have a dependency relationship within the fused operator node can be adjusted to achieve the adjustment of the peak memory usage of the operator nodes.
[0042] Optionally, a greedy algorithm is used to search backward with reference to the dynamic type to obtain supplementary fusion operator nodes, including: searching backward for single-input single-output or single-input multiple-output specified operator nodes to be fused that have a dependency relationship with the current operator node; obtaining the peak memory usage of the specified operator nodes to be fused and the backward search anchor point in the shared storage layer; and determining the supplementary fusion operator nodes using a greedy algorithm based on the peak memory usage and the storage capacity of the shared storage layer.
[0043] Specifically, after obtaining the updated computational graph through forward search in this embodiment, to avoid missing the search for fusion operator nodes, in this embodiment, the search will continue in the backward search manner. When performing the backward search, specifically, single-input multiple-output and multiple-input multiple-output specified operator nodes are used as the backward search anchor points. Here, the specified operator nodes can be candidate operator nodes that have not been fused or fusion operator nodes obtained through fusion. Then, search backward for single-input single-output or single-input multiple-output candidate operator nodes or fusion operator nodes to be fused that have a dependency relationship with the current operator node. And when performing the backward search, the greedy algorithm is still used. Therefore, the manner of the backward search is roughly the same as that of the forward search, and will not be elaborated in this embodiment. For example, the supplementary fusion operator node obtained through backward search is D3(CDE), and the updated computational graph is updated according to the supplementary fusion operator node to obtain the optimized computational graph. Therefore, the optimized computational graph includes two operator nodes, A2(AB) and D3(CDE). Of course, this is only an example in this embodiment, and the specific form of the obtained optimized computational graph is not limited.
[0044] In this embodiment, candidate operator nodes that meet the fusion conditions and the dynamic types of each candidate operator node are screened out from the computational graph, and the fusion operator nodes are obtained by searching for the candidate operator nodes according to the dynamic types. Since the candidate operator nodes located in the same fusion operator node only perform tensor transfer in the shared storage layer, the number of tensor data transfers between different storage layers is avoided, improving the data processing efficiency and computational performance of the computational graph.
[0045] Embodiment 2
[0046] Figure 3 The figure is a flowchart of a compilation optimization method for a computational graph provided in Embodiment 2 of the present invention. Based on the above embodiment, after obtaining the optimized computational graph by updating the computational graph according to the fusion operator nodes, it further includes: uniformly registering an operator interface for the fusion operator nodes included in the optimized computational graph, and configuring the attribute information of the fusion operator nodes through the operator interface. When an execution instruction for the optimized computational graph is received, the optimized computational graph is parsed and executed by calling the operator interface. As Figure 3 shown, the method includes:
[0047] Step S201: Screen the computation graph running on the multi-level storage chip to obtain candidate operator nodes.
[0048] Optionally, screening the computation graph running on the multi-level storage chip to obtain candidate operator nodes includes: obtaining the residence status of each operator node in the shared storage layer in the computation graph, where the residence status includes meeting the residence condition and not meeting the residence condition; querying the support status of the input or output of each operator node in the computation graph in the shared storage layer by calling the query interface, where the support status includes supported and not supported; and taking the operator nodes with the residence status of meeting the residence condition and the support status of supported as candidate operator nodes.
[0049] Optionally, obtaining the residence status of each operator node in the shared storage layer in the computation graph includes: obtaining the tensor size of each operator node in the computation graph and the storage capacity of the shared storage layer; determining whether the tensor size of each operator node is less than the storage capacity, and if so, determining that the residence status of the operator node in the shared storage layer meets the residence condition, otherwise, determining that the residence status of the operator node in the shared storage layer does not meet the residence condition.
[0050] Step S202: Obtain the dynamic types of the current candidate operator nodes.
[0051] Optionally, obtaining the dynamic types of the current candidate operator nodes includes: obtaining the number of non-constant input tensors of each current candidate operator node and taking the number of non-constant input tensors as the in-degree of the candidate operator node; obtaining the total number of output tensors of each current candidate operator node and taking the total number of output tensors as the out-degree of the candidate operator node; and determining the dynamic types of the current candidate operator nodes according to the in-degree and out-degree.
[0052] Step S203: Perform forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic types to obtain fused operator nodes, and update the computation graph according to the fused operator nodes to obtain an optimized computation graph.
[0053] Optionally, performing forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic types to obtain fused operator nodes, and updating the computation graph according to the fused operator nodes to obtain an optimized computation graph includes: taking a specified operator node in the computation graph as the forward search anchor point, and using the greedy algorithm to search forward and fuse with reference to the dynamic type to obtain at least one fused operator node, where the specified operator node includes a candidate operator node or a fused operator node; updating the computation graph according to the fused operator nodes to obtain an updated computation graph; taking the specified operator nodes with single input and multiple outputs and multiple inputs and multiple outputs in the updated computation graph as the backward search anchor points, and using the greedy algorithm to search backward with reference to the dynamic type to obtain supplementary fused operator nodes; and updating the updated computation graph according to the supplementary fused operator nodes to obtain an optimized computation graph.
[0054] Optionally, a greedy algorithm is used to search forward and fuse with reference to the dynamic type to obtain at least one fused operator node, including: searching forward for a to-be-fused specified operator node with a single-input single-output or multi-input single-output dependency relationship with the current specified operator node; obtaining the peak memory usage of the to-be-fused specified operator node and the forward search anchor point in the shared storage layer; and determining the fused operator node using a greedy algorithm based on the peak memory usage and the storage capacity of the shared storage layer.
[0055] Optionally, a greedy algorithm is used to determine the fused operator node based on the peak memory usage and the storage capacity of the shared storage layer, including: determining whether the peak memory usage exceeds the storage capacity of the shared storage layer; if so, adjusting the order of the to-be-fused specified operator nodes without a dependency relationship, and obtaining the fused operator node according to the adjustment result; otherwise, fusing the to-be-fused specified operator node with the forward search anchor point to obtain an intermediate fused operator node, updating the dynamic type of the intermediate fused operator node, using the intermediate fused operator node as a new forward search anchor point to increase the search range and re-search, and obtaining the fused operator node according to the result of the re-search.
[0056] Optionally, obtaining the fused operator node according to the adjustment result includes: determining whether there is an adjustment result where the peak memory usage does not exceed the storage capacity of the shared storage layer; if so, obtaining the operator node sequence corresponding to the adjustment result, and fusing the to-be-fused specified operator node and the anchor point according to the operator node sequence to obtain the fused operator node; otherwise, abandoning the fusion of the anchor point and the to-be-fused specified operator node.
[0057] Optionally, a greedy algorithm is used to search backward with reference to the dynamic type to obtain supplementary fused operator nodes, including: searching backward for a to-be-fused specified operator node with a single-input single-output or single-input multi-output dependency relationship with the current operator node; obtaining the peak memory usage of the to-be-fused specified operator node and the backward search anchor point in the shared storage layer; and determining the supplementary fused operator node using a greedy algorithm based on the peak memory usage and the storage capacity of the shared storage layer.
[0058] Step S204, uniformly register an operator interface for the fused operator nodes included in the optimized computation graph, and configure the attribute information of the fused operator nodes through the operator interface.
[0059] Specifically, in this embodiment, after obtaining the optimized computational graph through fusion optimization, when it is necessary to execute the optimized computational graph, it needs to be parsed first. Generally, for customized fusion, manual parsing is usually adopted. However, for the automatically generated fusion operator nodes in this application, neither the quantity, structure, memory operator type, nor the data type is fixed. Therefore, an automated parsing method needs to be adopted. In the embodiment of this application, in order to facilitate the implementation of subsequent automated parsing, an operator interface, such as AutoFusionOp, will be uniformly registered for all the fusion operator nodes included in the optimized computational graph. Of course, in this embodiment, only an example is given and the specific format of the operator interface is not limited. And the fusion operator nodes and ordinary operator nodes are at the same level for the framework-side scheduling and distribution. The upper-layer framework does not need to manage the inner-layer storage, that is, the shared storage layer L2, which simplifies the framework processing. In this embodiment, after the operator interface registration is completed, the attribute information of the fusion operator nodes will be configured through the operator interface. Among them, the attribute information includes the execution order, execution parameters, and the storage memory spaces of the input tensors and output tensors.
[0060] Among them, the execution order refers to the fusion order of each operator node sorted out during the fusion of the fusion operator nodes; since the complete operator expression is retained during the fusion, the execution parameters can be extracted during the parsing and integrated into the context parameter information of a single operator node, realizing the basis for calling a single operator node within the fusion operator node; if the input tensor of the operator comes from the previous operator inside the fusion, it is configured as the inner-layer storage managed inside AutoFusionOp; if it comes from outside the fusion, it is configured as the outer-layer storage uniformly managed by the framework side. In addition, if the output tensor of the operator is not the final output of the fusion operator, it is configured to be written back to the inner-layer storage managed inside AutoFusionOp; if it is the final output of the fusion operator, it is configured to be written back to the outer-layer storage uniformly managed by the framework side. Through the above configuration, it can be ensured that the upper-layer framework only manages the outer-layer storage, and the operator side and the underlying scheduler still manage the inner-layer storage, adapting the software stack at the lowest cost.
[0061] Step S205, when receiving an execution instruction for the optimized computational graph, parse and execute the optimized computational graph by calling the operator interface.
[0062] Optionally, parsing and executing the optimized computational graph by calling the operator interface includes: disassembling the fusion operator nodes in the optimized computational graph by calling the operator interface to obtain candidate operator nodes, and obtaining the attribute information of the candidate operator nodes by calling the operator interface; executing the candidate operator nodes according to the execution parameters in the execution order, and saving the execution results of each candidate operator node in the corresponding memory space.
[0063] Specifically, in this embodiment, when an execution instruction for the optimized computation graph is received, the operator interface is called. All the fused operator nodes obtained through automatic fusion will enter the operator interface to complete the computation. Moreover, in the operator interface, the fused operator nodes will be disassembled to obtain candidate operator nodes. Additionally, the attribute information of each pre-configured candidate operator node will be obtained through the operator interface, and the candidate operator nodes will be executed based on the execution parameters in the execution order. For the execution results of each candidate operator node, the input tensors and output tensors will be correspondingly saved in the memory space configured in the operator interface.
[0064] It is worth mentioning that for the operator nodes that are not fused, each time they are called, a dispatch operation needs to be executed, which brings corresponding fixed overheads. However, in this application, by calling the operator interface to parse the fused operator nodes, only one dispatch operation needs to be executed to achieve the dispatch of multiple operator nodes, thereby eliminating the excessive fixed overheads caused by multiple dispatches. This has a significant optimization effect for operator nodes with short execution times. Additionally, by performing a lifecycle analysis on the temporary result data stored in the inner-layer memory, the application and release of the inner-layer memory can be overall grasped, and the inner-layer memory space can be fully utilized.
[0065] In this embodiment, candidate operator nodes that meet the fusion conditions and the dynamic types of each candidate operator node are screened out from the computation graph, and the fused operator nodes are obtained by searching for the candidate operator nodes according to the dynamic types. Since the candidate operator nodes located in the same fused operator node only perform tensor transfers in the shared memory layer, the number of tensor data transfers between different memory layers is avoided, thereby improving the data processing efficiency and computing performance of the computation graph.
[0066] Embodiment Three
[0067] Figure 4 FIG. is a schematic structural diagram of a computation graph compilation optimization device provided in Embodiment Three of the present invention. As Figure 4 shown, the device includes: a candidate operator node screening module 310, a dynamic type obtaining module 320, and an optimized computation graph obtaining module 330.
[0068] Among them, the candidate operator node screening module 310 is configured to screen the computation graph running on a multi-level storage chip to obtain candidate operator nodes;
[0069] The dynamic type obtaining module 320 is configured to obtain the dynamic types of the current candidate operator nodes, where the dynamic types include single-input single-output, single-input multi-output, multi-input single-output, and multi-input multi-output;
[0070] The optimized computation graph acquisition module 330 is used to perform forward search and backward search on candidate operator nodes in the computation graph according to the dynamic type to obtain fused operator nodes, and update the computation graph according to the fused operator nodes to obtain the optimized computation graph.
[0071] Among them, all candidate operator nodes in the fused operator nodes move non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip.
[0072] Optionally, the candidate operator node screening module is used to obtain the residence status of each operator node in the shared storage layer, where the residence status includes meeting the residence condition and not meeting the residence condition;
[0073] Query the support status of the input or output of each operator node in the shared storage layer by calling the query interface, where the support status includes supported and not supported;
[0074] Use the operator nodes with the residence status of meeting the residence condition and the support status of supported as candidate operator nodes.
[0075] Optionally, the candidate operator node screening module is also used to obtain the tensor size of each operator node in the computation graph and the storage capacity of the shared storage layer;
[0076] Judge whether the tensor size of each operator node is less than the storage capacity. If so, determine that the residence status of the operator node in the shared storage layer meets the residence condition.
[0077] Otherwise, determine that the residence status of the operator node in the shared storage layer does not meet the residence condition.
[0078] Optionally, the dynamic type acquisition module is used to obtain the current number of non-constant input tensors of each current candidate operator node, and use the number of non-constant input tensors as the in-degree of the candidate operator node;
[0079] Obtain the current total number of output tensors of each current candidate operator node, and use the total number of output tensors as the out-degree of the candidate operator node;
[0080] Determine the dynamic type of each current candidate operator node according to the in-degree and out-degree.
[0081] Optionally, the optimized computation graph acquisition module includes: a forward search unit, which is used to use the specified operator node in the computation graph as the forward search anchor point, and adopt the greedy algorithm to search forward and fuse according to the dynamic type to obtain at least one fused operator node, where the specified operator node includes a candidate operator node or a fused operator node;
[0082] The updated computation graph acquisition unit is used to update the computation graph according to the fused operator nodes to obtain the updated computation graph;
[0083] A backward search unit, which is used to use the specified operator nodes with single input and multiple outputs and multiple inputs and multiple outputs in the updated computational graph as backward search anchor points, and uses a greedy algorithm to search backward with reference to the dynamic type to obtain supplementary fusion operator nodes;
[0084] An optimized computational graph acquisition unit, which is used to update the updated computational graph according to the supplementary fusion operator nodes to obtain an optimized computational graph.
[0085] Optionally, a forward search unit, which is used to search forward for the specified operator nodes to be fused with single input and single output or multiple inputs and single output that have a dependency relationship with the current specified operator node;
[0086] Obtain the peak memory usage of the specified operator node to be fused and the forward search anchor point in the shared storage layer;
[0087] Determine the fusion operator nodes according to the peak memory usage and the storage capacity of the shared storage layer by using a greedy algorithm.
[0088] Optionally, the forward search unit is further used to determine whether the peak memory usage exceeds the storage capacity of the shared storage layer. If so, perform a sequential adjustment on the specified operator nodes to be fused that have no dependency relationship, and obtain the fusion operator nodes according to the adjustment result.
[0089] Otherwise, fuse the specified operator node to be fused with the forward search anchor point to obtain an intermediate fusion operator node, update the dynamic type of the intermediate fusion operator node, use the intermediate fusion operator node as a new forward search anchor point to increase the search range and perform a search again, and obtain the fusion operator nodes according to the results of the re-search.
[0090] Optionally, the forward search unit is further used to determine whether there is an adjustment result where the peak memory usage does not exceed the storage capacity of the shared storage layer. If so, obtain the operator node sequence corresponding to the adjustment result, and fuse the specified operator node to be fused and the anchor point according to the operator node sequence to obtain the fusion operator nodes.
[0091] Otherwise, abandon the fusion of the anchor point and the specified operator node to be fused.
[0092] Optionally, the backward search unit is used to search backward for the specified operator nodes to be fused with single input and single output or single input and multiple outputs that have a dependency relationship with the current operator node;
[0093] Obtain the peak memory usage of the specified operator node to be fused and the backward search anchor point in the shared storage layer;
[0094] Determine the supplementary fusion operator nodes according to the peak memory usage and the storage capacity of the shared storage layer by using a greedy algorithm.
[0095] Optionally, the device further includes an operator interface registration module, configured to uniformly register an operator interface for the fused operator nodes included in the optimized computation graph, and configure the attribute information of the fused operator nodes through the operator interface, where the attribute information includes the execution order, execution parameters, and the storage memory spaces of the input tensor and the output tensor;
[0096] A parsing and execution module, configured to, when receiving an execution instruction for the optimized computation graph, parse and execute the optimized computation graph by calling the operator interface.
[0097] Optionally, the parsing and execution module is further configured to disassemble the fused operator nodes in the optimized computation graph by calling the operator interface to obtain candidate operator nodes,
[0098] obtain the attribute information of the candidate operator nodes by calling the operator interface;
[0099] execute the candidate operator nodes according to the execution parameters in the execution order, and save the execution results of the candidate operator nodes in the corresponding memory spaces.
[0100] The computation graph compilation and optimization device provided by the embodiments of the present invention can execute the computation graph compilation and optimization method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0101] Embodiment 4
[0102] Figure 5 FIG. is a schematic structural diagram of a computer device provided by Embodiment 4 of the present invention. As Figure 5 shown, the computer device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the computer device may be one or more, Figure 5 taking one processor 610 as an example; the processor 610, the memory 620, the input device 630, and the output device 640 in the computer device may be connected through a bus or other means, Figure 5 taking the connection through a bus as an example.
[0103] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the computation graph compilation and optimization method in the embodiments of the present invention. The processor 610 runs the software programs, instructions, and modules stored in the memory 620, thereby executing various functional applications and data processing of the computer device, that is, implementing the above-mentioned computation graph compilation and optimization method, including:
[0104] screen the computation graph running on the multi-level storage chips to obtain candidate operator nodes;
[0105] Obtain the dynamic types of current candidate operator nodes, where the dynamic types include single-input single-output, single-input multiple-output, multiple-input single-output, and multiple-input multiple-output;
[0106] Perform forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic types to obtain fused operator nodes, and update the computation graph according to the fused operator nodes to obtain an optimized computation graph,
[0107] Among them, all candidate operator nodes in the fused operator nodes perform the transfer of non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip.
[0108] The memory 620 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 620 may further include a memory remotely set relative to the processor 610, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0109] The input device 630 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device. The output device 640 may include a display device such as a display screen.
[0110] Embodiment Five
[0111] Embodiment Five of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a compilation optimization method of a computation graph when executed by a computer processor, including:
[0112] Screen the computation graph running on the multi-level storage chip to obtain candidate operator nodes;
[0113] Obtain the dynamic types of current candidate operator nodes, where the dynamic types include single-input single-output, single-input multiple-output, multiple-input single-output, and multiple-input multiple-output;
[0114] Perform forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic types to obtain fused operator nodes, and update the computation graph according to the fused operator nodes to obtain an optimized computation graph,
[0115] Among them, all candidate operator nodes in the fused operator nodes perform the transfer of non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip.
[0116] Certainly, for the storage medium containing computer-executable instructions provided by the embodiments of the present invention, the computer-executable instructions are not limited to the method operations as described above, and can also execute the related operations in the compilation optimization method of the computational graph provided by any embodiment of the present invention.
[0117] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present invention.
[0118] It should be noted that in the embodiments of the above compilation optimization device of the computational graph, the included units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0119] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for compiling and optimizing a computational graph, characterized in that: include: Screen the computational graph running on the multi-level storage chip to obtain candidate operator nodes; Obtaining the dynamic type of each of the current candidate operator nodes, wherein the dynamic type includes single-input single-output, single-input multiple-output, multiple-input single-output, and multiple-input multiple-output; According to the dynamic type, forward search and backward search are performed on the candidate operator nodes in the calculation graph to obtain a fusion operator node, and according to the fusion operator node, the calculation graph is updated to obtain an optimized calculation graph, Wherein, all candidate operator nodes in the fusion operator node carry out the transfer of non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip; The step of performing forward search and backward search on the candidate operator nodes in the computation graph according to the dynamic type to obtain a fused operator node, and updating the computation graph according to the fused operator node to obtain an optimized computation graph includes: Taking the designated operator node in the computation graph as a forward search anchor point, a greedy algorithm is used to forward search and fuse to obtain at least one fused operator node with reference to the dynamic type, wherein the designated operator node includes a candidate operator node or a fused operator node; Updating the computation graph according to the fusion operator node to obtain an updated computation graph; Using the designated SIMO and MIMO operator nodes in the update calculation graph as backward search anchor points, and using the greedy algorithm to perform backward search with reference to the dynamic type to obtain a supplementary fusion operator node; The update calculation graph is updated according to the supplementary fusion operator node to obtain the optimized calculation graph.
2. The method according to claim 1, characterized in that The step of screening the computation graph running on the multi-level storage chip to obtain candidate operator nodes includes: Obtaining the residency status of each operator node in the computation graph in the shared storage layer, wherein the residency status includes having a residency condition and not having a residency condition; Query the support status of the input or output of each operator node in the computation graph in the shared storage layer by calling the query interface, wherein the support status includes supported and unsupported; The operator node whose resident status is satisfied with the resident condition and whose support status is supported is taken as the candidate operator node.
3. The method according to claim 2, characterized in that The obtaining the resident status of each operator node in the computation graph in the shared storage layer includes: Obtaining the tensor size of each operator node in the computation graph and the storage capacity of the shared storage layer; Determine whether the tensor size of each operator node is less than the storage capacity, and if so, determine that the resident state of the operator node in the shared storage layer is the resident condition. Otherwise, it is determined that the residency state of the operator node in the shared storage layer does not meet the residency condition.
4. The method according to claim 1, characterized in that: The obtaining of the dynamic type of each of the current candidate operator nodes includes: Obtain the current number of non-constant input tensors of each of the candidate operator nodes, and use the number of non-constant input tensors as the in-degree of the candidate operator node; Obtain the current total number of output tensors of each of the candidate operator nodes, and use the total number of output tensors as the out-degree of the candidate operator node; The dynamic type of each current candidate operator node is determined according to the in-degree and the out-degree.
5. The method according to claim 1, characterized in that The step of using a greedy algorithm to search forward with reference to the dynamic type and fusing to obtain at least one fusion operator node includes: Search forward for a single-input single-output or multi-input single-output designated operator node to be fused that has a dependency relationship with the current designated operator node; Obtaining the peak memory usage of the designated operator node to be fused and the forward search anchor point in the shared storage layer; The fusion operator node is determined by using the greedy algorithm according to the memory usage peak and the storage capacity of the shared storage layer.
6. The method according to claim 5, characterized in that The step of determining the fusion operator node by using the greedy algorithm according to the memory usage peak value and the storage capacity of the shared storage layer includes: Determine whether the peak memory usage exceeds the storage capacity of the shared storage layer. If so, sequentially adjust the designated operator nodes to be fused that do not have a dependency relationship, and obtain the fused operator node according to the adjustment result. Otherwise, the designated operator node to be fused is fused with the forward search anchor point to obtain an intermediate fusion operator node, the dynamic type of the intermediate fusion operator node is updated, the intermediate fusion operator node is used as a new forward search anchor point to increase the search range and search again, and the fusion operator node is obtained according to the result of the re-search.
7. The method according to claim 6, characterized in that The obtaining the fusion operator node according to the adjustment result includes: Determine whether there is an adjustment result in which the peak value of memory usage does not exceed the storage capacity of the shared storage layer. If so, obtain the operator node sequence corresponding to the adjustment result, and fuse the designated operator node to be fused and the anchor point according to the operator node sequence to obtain the fused operator node. Otherwise, the fusion of the anchor point and the designated operator node to be fused is abandoned.
8. The method according to claim 1, characterized in that The step of using a greedy algorithm to search backward with reference to the dynamic type to obtain a supplementary fusion operator node includes: Search backward for a single-input single-output or single-input multiple-output designated operator node to be fused that has a dependency relationship with the current operator node; Obtaining the peak memory usage of the designated operator node to be fused and the backward search anchor point in the shared storage layer; The supplementary fusion operator node is determined by using the greedy algorithm according to the memory usage peak and the storage capacity of the shared storage layer.
9. The method according to claim 1, characterized in that: After the calculation graph is updated according to the fusion operator node to obtain the optimized calculation graph, the method further includes: Registering an operator interface uniformly for the fusion operator nodes included in the optimization calculation graph, and configuring the attribute information of the fusion operator nodes through the operator interface, wherein the attribute information includes execution order, execution parameters, and storage memory space of input tensors and output tensors; When an execution instruction for the optimization calculation graph is received, the optimization calculation graph is parsed and executed by calling the operator interface.
10. The method according to claim 9, characterized in that The parsing and executing the optimization calculation graph by calling the operator interface includes: By calling the operator interface, the fusion operator node in the optimization calculation graph is disassembled to obtain candidate operator nodes, Acquire the attribute information of the candidate operator node by calling the operator interface; The candidate operator nodes are executed according to the execution parameters based on the execution order, and the execution results of each candidate operator node are saved in a corresponding memory space.
11. A computing graph compilation optimization device, characterized in that: include: A candidate operator node screening module is used to screen the computation graph running on the multi-level storage chip to obtain candidate operator nodes; A dynamic type acquisition module, used to acquire the dynamic type of each of the current candidate operator nodes, wherein the dynamic type includes single-input single-output, single-input multiple-output, multiple-input single-output and multiple-input multiple-output; An optimized calculation graph acquisition module is used to perform forward search and backward search on the candidate operator nodes in the calculation graph according to the dynamic type to obtain a fusion operator node, and update the calculation graph according to the fusion operator node to obtain an optimized calculation graph, Wherein, all candidate operator nodes in the fusion operator node carry out the transfer of non-constant input tensors and output tensors in the shared storage layer of the multi-level storage chip; The optimized calculation graph acquisition module is further used to use the designated operator node in the calculation graph as a forward search anchor point, adopt a greedy algorithm to forward search and fuse to obtain at least one fusion operator node with reference to the dynamic type, wherein the designated operator node includes a candidate operator node or a fusion operator node; Updating the computation graph according to the fusion operator node to obtain an updated computation graph; Using the designated SIMO and MIMO operator nodes in the update calculation graph as backward search anchor points, and using the greedy algorithm to perform backward search with reference to the dynamic type to obtain a supplementary fusion operator node; The update calculation graph is updated according to the supplementary fusion operator node to obtain the optimized calculation graph.
12. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 10 is implemented.
13. A computer executable instruction storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Automatic operator fusion method of computational graph and related product
CN115756478A