A method, device, storage medium and program product for computing graph optimization

By splitting and merging the calculation diagrams of the encoding modules on the artificial intelligence chip, generating and compiling them into kernel functions, the problem of poor optimization of the encoding module calculation diagrams in the existing technology is solved, and the execution performance is improved.

CN119337046BActive Publication Date: 2025-07-08SHANGHAI BIREN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411874219.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-08
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The common graph optimization strategy in the existing deep learning framework is difficult to effectively optimize the computational graph of the encoding module, resulting in poor execution performance.

Method used

Based on the index resource amount of artificial intelligence chip, the calculation diagram of the encoding module is divided into multiple calculation subgraphs, and the subgraph is merged and merged according to the subgraph merging granularity, and merged it, and compiled into a merged kernel function for execution.

Benefits of technology

By reducing the number of merged kernel functions, the optimization effect of the calculation graph is improved and the execution performance of the encoding module is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119337046B_ABST
    Figure CN119337046B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a computational graph optimization method, device, storage medium, and program product, which relate to the field of artificial intelligence technology. The method includes: splitting the computational graph of an encoding module into multiple computational subgraphs based on a preset template; obtaining a subgraph merging granularity based on the resource amount of the index resources on an artificial intelligence chip; and then performing assembly merging on the multiple computational subgraphs according to the subgraph merging granularity to obtain at least one merged subgraph. That is, when fully considering the index resources on the artificial intelligence chip, the computational graph is split into as few merged subgraphs as possible. Therefore, when each merged subgraph is compiled into a corresponding merged kernel function and executed, the number of merged kernel functions is greatly reduced, thereby reducing the time overhead of launching kernel functions. Secondly, the present application not only considers the operators included in the computational graph itself, but also combines the resource situation of the underlying chip that actually executes the computational graph, enabling the computational graph to achieve maximum optimization and improving the optimization effect of the computational graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a computational graph optimization method, device, storage medium, and program product. Background Art

[0002] In the field of deep learning, the encoding module (Encoder) has been widely used in many models. For example, pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers). In the era of large models, the importance of the encoding module has become even more prominent. Therefore, optimizing the computational graph of the encoding module to improve its execution performance is particularly important.

[0003] Under the related technology, the computational graph is optimized through general graph optimization strategies in the deep learning framework. Considering generality, these optimization strategies often target some common scenarios in deep learning models and have no association with the underlying chips that actually execute model inference. Therefore, when using this method to optimize the computational graph of the encoding module, it is difficult to achieve a better optimization effect. Summary of the Invention

[0004] Embodiments of the present application provide a computational graph optimization method, device, storage medium, and program product, which are used to improve the optimization effect of the computational graph of the encoding module, thereby improving the execution performance of the encoding module.

[0005] On the one hand, embodiments of the present application provide a computational graph optimization method, including:

[0006] Segmenting the computational graph of the encoding module into multiple computational subgraphs based on a preset template;

[0007] Obtaining a subgraph merging granularity based on the resource amount of the index resources on the artificial intelligence chip;

[0008] Assembling and merging the multiple computational subgraphs according to the subgraph merging granularity to obtain at least one merged subgraph;

[0009] Compiling each merged subgraph into a corresponding merged kernel function and executing the merged kernel function corresponding to each merged subgraph to obtain the computational result of the encoding module.

[0010] On the one hand, embodiments of the present application provide a computational graph optimization device, including:

[0011] A segmentation module, configured to segment the computational graph of the encoding module into multiple computational subgraphs based on a preset template;

[0012] A merging module, configured to obtain a sub - graph merging granularity based on the resource amount of index resources on an artificial intelligence chip; and perform assembly merging on the multiple computing sub - graphs according to the sub - graph merging granularity to obtain at least one merged sub - graph.

[0013] An execution module, configured to compile each merged sub - graph into a corresponding merged kernel function and execute the merged kernel function corresponding to each merged sub - graph to obtain the calculation result of the encoding module.

[0014] Optionally, the index resources are used to: index video memory objects for input tensors or output tensors;

[0015] Specifically, the merging module is configured to:

[0016] Based on the resource amount of the index resources, determine an upper limit value of the total number of input tensors and output tensors included in a merged sub - graph;

[0017] Use the upper limit value as the sub - graph merging granularity.

[0018] Optionally, specifically, the merging module is configured to:

[0019] Perform assembly merging on the multiple computing sub - graphs in sequence according to the sub - graph merging granularity and the execution order of the multiple computing sub - graphs to obtain at least one merged sub - graph.

[0020] Optionally, the merging module is further configured to:

[0021] After performing assembly merging on the multiple computing sub - graphs according to the sub - graph merging granularity to obtain at least one merged sub - graph, for each merged sub - graph, respectively perform the following operations:

[0022] Traverse each intermediate tensor of a merged sub - graph, and for each traversed intermediate tensor, allocate cache resources in the artificial intelligence chip for the intermediate tensor, where the intermediate tensor is a tensor passed between computing sub - graphs in the merged sub - graph.

[0023] Optionally, specifically, the merging module is configured to:

[0024] For each traversed intermediate tensor, when it is determined that the storage space occupied by the intermediate tensor is less than a preset threshold, allocate the cache resources for the intermediate tensor.

[0025] Optionally, the cache resources allocated for an intermediate tensor do not overlap with the cache resources allocated for other intermediate tensors between the computing sub - graphs on which the intermediate tensor depends.

[0026] Optionally, the merging module is further configured to:

[0027] Before combining and assembling the multiple computing sub - graphs according to the sub - graph combination granularity to obtain at least one combined sub - graph, for each computing sub - graph, if the computing sub - graph includes at least one operator combination for layout adjustment, the following sub - graph optimization operations are performed on the computing sub - graph:

[0028] Replace the at least one operator combination for layout adjustment included in the computing sub - graph with the output address corresponding to each of the at least one operator combination for layout adjustment.

[0029] On the one hand, an embodiment of the present application provides a computer device, including a memory, an artificial intelligence chip, and a computer program stored on the memory and executable on the artificial intelligence chip. When the artificial intelligence chip executes the computer program, the steps of the above - mentioned computing graph optimization method are implemented.

[0030] On the one hand, an embodiment of the present application provides a computer - readable storage medium, which stores a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is enabled to execute the steps of the above - mentioned computing graph optimization method.

[0031] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer - readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is enabled to execute the steps of the above - mentioned computing graph optimization method.

[0032] In the embodiment of the present application, the computing graph of the encoding module is segmented into multiple computing sub - graphs based on a preset template; then, based on the resource amount of the index resources on the artificial intelligence chip, the sub - graph combination granularity is obtained; and then, according to the sub - graph combination granularity, the multiple computing sub - graphs are combined and assembled to obtain at least one combined sub - graph. That is, when fully considering the index resources on the artificial intelligence chip, the computing graph is segmented into as few combined sub - graphs as possible. Therefore, when each combined sub - graph is compiled into a corresponding combined kernel function and executed, the number of combined kernel functions is greatly reduced, thereby reducing the time overhead of launching kernel functions. Secondly, the computing graph optimization method of the present application not only considers the operators included in the computing graph itself, but also combines the resource situation of the underlying chip that actually executes the computing graph, so that the computing graph is optimized to the maximum extent and the optimization effect of the computing graph is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained without creative efforts based on these drawings.

[0034] Figure 1 Schematic diagram of the structure of an artificial intelligence chip provided by an embodiment of the present application;

[0035] Figure 2 Schematic flowchart of a computational graph optimization method provided by an embodiment of the present application;

[0036] Figure 3 Computational graph of the forward encoding of BERT provided by an embodiment of the present application;

[0037] Figure 4 Schematic diagram of the result of splitting a computational graph provided by an embodiment of the present application Figure 1 ;

[0038] Figure 5 Computational graph of the backward encoding of BERT provided by an embodiment of the present application;

[0039] Figure 6 Schematic diagram of the result of splitting a computational graph provided by an embodiment of the present application Figure 2 ;

[0040] Figure 7 Schematic flowchart of a sub - graph optimization method provided by an embodiment of the present application Figure 1 ;

[0041] Figure 8 Schematic flowchart of a sub - graph optimization method provided by an embodiment of the present application Figure 2 ;

[0042] Figure 9 Schematic flowchart of a cache resource allocation method provided by an embodiment of the present application Figure 1 ;

[0043] Figure 10 Schematic flowchart of a cache resource allocation method provided by an embodiment of the present application Figure 2 ;

[0044] Figure 11 Schematic diagram of the structure of a computational graph optimization device provided by an embodiment of the present application;

[0045] Figure 12 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0046] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] Reference Figure 1 FIG. 1 is a structural diagram of an artificial intelligence chip applicable to an embodiment of the present application. The artificial intelligence chip 100 at least includes: a video memory 101 and a plurality of execution units 102. Each execution unit 102 includes: an on-chip cache 103 and a register 104.

[0048] The video memory 101 can be a High Bandwidth Memory (HBM) or other types of memories. The on-chip cache 103 is a temporary memory with a smaller capacity than the video memory 101 but a faster data exchange speed than the video memory 101. The on-chip cache 103 can be a Gemm Main Buffer (GMB).

[0049] Compared with the on-chip cache 103, the register 104 has a smaller capacity than the on-chip cache 103 but a faster data exchange speed than the on-chip cache 103. The register 104 can be a Thread-Local Register (TLR).

[0050] In addition to the above structure, the artificial intelligence chip 100 in the present application may further include other structures, which are not specifically limited in the present application.

[0051] The artificial intelligence chip 100 can be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a General-purpose computing on graphics processing units (GPGPU), a Domain Specific Architecture (DSA), etc.

[0052] In the field of deep learning, the encoding module has been widely used in many models. Especially in sequence modeling and natural language processing, the encoding module in the pre-trained language model BERT is a core component of the Transformer architecture, which represents the logic composed of a series of operators. In the era of large models, the importance of the encoding module has become even more prominent. Therefore, it is particularly important to optimize the computational graph of the encoding module to improve the execution performance of the encoding module.

[0053] In the related art, the computational graph is optimized through general graph optimization strategies in deep learning frameworks. Considering generality, these optimization strategies often target some common scenarios in deep learning models and have no association with the underlying chips that actually execute model inference. Therefore, when using this method to optimize the computational graph of the encoding module, it is difficult to achieve a better optimization effect.

[0054] In view of this, based on Figure 1 the architecture diagram of the artificial intelligence chip shown, a computational graph optimization method is provided, mainly involving the optimization of the computational graph of the encoding module. In this application, the encoding module is used to convert input data into a vector or feature representation of a fixed size. The encoding module can be applied to various scenarios. For example, image processing scenarios, speech processing scenarios, text processing scenarios, etc. In different application scenarios, the physical meaning of the input data of the encoding module can be different.

[0055] For example, in the speech processing scenario, the input data of the encoding module can be speech data used in tasks such as speech enhancement, speech recognition, and speech synthesis. For example, in the image processing scenario, the input data of the encoding module can be image data used in tasks such as image preprocessing, image segmentation, and object detection. For example, in the text processing scenario, the input data of the encoding module can be text data used in tasks such as text generation and text recognition.

[0056] The following specifically introduces the process of the computational graph optimization method. The process of this method is executed by a computer device, and the computer device includes Figure 1 the artificial intelligence chip shown, and includes the following steps:

[0057] Step 201, split the computational graph of the encoding module into multiple computational subgraphs based on a preset template.

[0058] Specifically, multiple templates (patterns) are preset in advance. The templates can be some complex fused operators, such as the fused operator of a convolution operator and a bitwise calculation operator, the fused operator of a matrix multiplication operator and a bitwise calculation operator, etc.; they can also be single operators, such as a normalization operator, a reduction operator, etc.; they can also be some special functions, such as a loss function.

[0059] Adopt the method of template matching (pattern match) to automatically split the computational graph of the encoding module into multiple computational subgraphs. In specific implementation, the computational graph is matched with each preset template. When the computational graph contains content that matches the template, the matched content is used as a computational subgraph; for the content in the computational graph that does not match any template, each original operator in the content is used as a computational subgraph.

[0060] In this way, each computational sub-graph includes at least one original operator in the computational graph. At least one original operator included in a computational sub-graph forms a fused operator. The number of computational sub-graphs obtained after switching is less than the number of original operators included in the computational graph.

[0061] In some embodiments, the computational graph of the encoding module is: the computational graph of the forward encoding module or the computational graph of the reverse encoding module. Among them, the structure of the computational graph of the forward encoding module is different from the structure of the computational graph of the reverse encoding module.

[0062] For example, referring to Figure 3 , which is the computational graph of the forward encoding module of BERT provided by the embodiments of the present application. The original operators included in this computational graph are respectively: Forward Layer Normalization (LayerNormFwd) operator, Matrix Multiplication (Matmul) operator, PointWise operator, Reshape operator, Permute operator, Forward Softmax (SoftmaxFwd) operator, etc.

[0063] The process of performing forward encoding according to the computational graph is as follows:

[0064] Input the input tensor into the Forward Layer Normalization operator 1 for calculation to obtain the layer normalization result 1. Among them, the input tensor can be the input data of the model or the output result of the previous module in the forward encoding module of the model.

[0065] Use the Matrix Multiplication operator 1, PointWise operator 1, Reshape operator 1, and Permute operator 1 to calculate the layer normalization result 1 in sequence to obtain the Permute result 1.

[0066] Use the Matrix Multiplication operator 2, PointWise operator 2, Reshape operator 2, and Permute operator 2 to calculate the layer normalization result 1 in sequence to obtain the Permute result 2.

[0067] Use the Matrix Multiplication operator 3, PointWise operator 3, Reshape operator 3, and Permute operator 3 to calculate the layer normalization result 1 in sequence to obtain the Permute result 3.

[0068] Input the Permute result 1 and the Permute result 2 into the Matrix Multiplication operator 4 for calculation to obtain the forward matrix multiplication result 1. Use the PointWise operator 4, PointWise operator 5, and Forward Softmax operator to calculate the forward matrix multiplication result 1 in sequence to obtain the normalization result.

[0069] Input the normalization result and the Permute result 3 into the Matrix Multiplication operator 5 for calculation to obtain the forward matrix multiplication result 2. Use the Permute operator 4, Reshape operator 4, Matrix Multiplication operator 6, and PointWise operator 6 to calculate the forward matrix multiplication result 2 in sequence to obtain the PointWise calculation result 1.

[0070] Input the layer normalization result 1 and the point-to-point calculation result 1 into the point-to-point operator 7 for calculation to obtain the point-to-point calculation result 2. Input the point-to-point calculation result 2 into the forward layer normalization operator 2 for calculation to obtain the layer normalization result 2.

[0071] Use the matrix multiplication operator 7, point-to-point operator 8, point-to-point operator 9, matrix multiplication operator 8, and point-to-point operator 10 to calculate the layer normalization result 2 in sequence to obtain the point-to-point calculation result 3.

[0072] Input the layer normalization result 2 and the point-to-point calculation result 3 into the point-to-point operator 11 for calculation to obtain the output tensor.

[0073] It is set that there are 3 templates preset, namely template 1, template 2, and template 3. Among them, template 1 is a fused operator composed of a matrix multiplication operator, a point-to-point operator, a deformation operator, and a reordering operator; template 2 is a fused operator composed of a matrix multiplication operator, a point-to-point operator, and a point-to-point operator; template 3 is a fused operator composed of a matrix multiplication operator, a reordering operator, and a deformation operator.

[0074] Adopt the method of template matching to Figure 3 The computational graph of is sliced into 11 computational subgraphs (i.e., forward subgraphs), see Figure 4 , which are respectively: forward sub Figure 1 to forward sub Figure 11 .

[0075] In specific implementation, since the forward layer normalization operator 1 in the computational graph does not match any of the above 3 templates, the forward layer normalization operator 1 is separately used as the forward sub Figure 1 .

[0076] Since the matrix multiplication operator 1, point-to-point operator 1, deformation operator 1, and reordering operator 1 in the computational graph match template 1, the matrix multiplication operator 1, point-to-point operator 1, deformation operator 1, and reordering operator 1 in the computational graph are divided into the forward sub Figure 2 .

[0077] And so on, other forward subgraphs can be obtained in the same way. The specific content of each of the other forward subgraphs is as follows:

[0078] Forward sub Figure 3 includes: matrix multiplication operator 2, point-to-point operator 2, deformation operator 2, reordering operator 2;

[0079] Forward sub Figure 4 includes: matrix multiplication operator 3, point-to-point operator 3, deformation operator 3, reordering operator 3;

[0080] Forward sub Figure 5 includes: matrix multiplication operator 4, point-to-point operator 4, point-to-point operator 5;

[0081] Forward sub Figure 6 including: forward normalization operator;

[0082] Forward sub Figure 7 including: matrix multiplication operator 5, reordering operator 4, deformation operator 4;

[0083] Forward sub Figure 8 including: matrix multiplication operator 6, point-to-point operator 6, point-to-point operator 7;

[0084] Forward sub Figure 9 including: forward layer normalization operator 2;

[0085] Forward sub Figure 10 including: matrix multiplication operator 7, point-to-point operator 8, point-to-point operator 9;

[0086] Forward sub Figure 11 including: matrix multiplication operator 8, point-to-point operator 10, point-to-point operator 11.

[0087] It can be seen from Figure 3 and Figure 4 that Figure 3 the computational graph contains 30 operators, and after slicing the computational graph shown in Figure 3 , 11 computational subgraphs are obtained, that is, the number of computational subgraphs contained in the computational graph is less than the number of operators contained in the computational graph.

[0088] For example, referring to Figure 5 , it is the computational graph of the reverse encoding module of BERT provided by the embodiment of the present application. The original operators contained in this computational graph are respectively: reverse layer normalization (LayerNormBwd) operator, reduction (Reduce) operator, matrix multiplication (Matmul) operator, point-to-point (PointWise) operator, deformation (Reshape) operator, reordering (Permute) operator, reverse normalization (SoftmaxFwd) operator, etc.

[0089] The process of performing reverse encoding according to the computational graph is as follows:

[0090] Use the reverse layer normalization operator 1 to calculate the gradient information to obtain the layer normalization result 3.

[0091] Use the reduction operator 1 to calculate the layer normalization result 3 to obtain the reduction result 1. Use the matrix multiplication operator 9 to calculate the layer normalization result 3 to obtain the reverse matrix multiplication result 1. Use the matrix multiplication operator 10 and the point-to-point operator 12 to calculate the layer normalization result 3 to obtain the point-to-point calculation result 4.

[0092] Use reduction operator 2 to calculate the point-to-point calculation result 4 to obtain reduction result 2. Use matrix multiplication operator 11 to calculate the point-to-point calculation result 4 to obtain reverse matrix multiplication result 2. Use matrix multiplication operator 12 and point-to-point operator 13 to calculate the point-to-point calculation result 4 in sequence to obtain point-to-point calculation result 5.

[0093] Use reverse layer normalization operator 2 to calculate the point-to-point calculation result 5 to obtain layer normalization result 4.

[0094] Use reduction operator 3 to calculate the layer normalization result 4 to obtain reduction result 3. Use matrix multiplication operator 13 to calculate the layer normalization result 4 to obtain reverse matrix multiplication result 3. Use matrix multiplication operator 14, deformation operator 5, and reordering operator 5 to calculate the layer normalization result 4 to obtain reordering result 4.

[0095] Use matrix multiplication operator 15, reverse normalization operator, and point-to-point operator 14 to calculate the reordering result 4 in sequence to obtain point-to-point calculation result 6.

[0096] Use matrix multiplication operator 16, reordering operator 6, and deformation operator 6 to calculate the reordering result 4 in sequence to obtain deformation result 1.

[0097] Use reduction operator 4 to calculate the deformation result 1 to obtain reduction result 4. Use matrix multiplication operator 17 to calculate the deformation result 1 to obtain reverse matrix multiplication result 4. Use matrix multiplication operator 18 to calculate the deformation result 1 to obtain reverse matrix multiplication result 5.

[0098] Use matrix multiplication operator 19, reordering operator 7, and deformation operator 7 to calculate the point-to-point calculation result 6 in sequence to obtain deformation result 2.

[0099] Use reduction operator 5 to calculate the deformation result 2 to obtain reduction result 5. Use matrix multiplication operator 20 to calculate the deformation result 2 to obtain reverse matrix multiplication result 6. Use matrix multiplication operator 21 to calculate the deformation result 2 to obtain reverse matrix multiplication result 7.

[0100] Use matrix multiplication operator 22, reordering operator 8, and deformation operator 8 to calculate the point-to-point calculation result 6 in sequence to obtain deformation result 3.

[0101] Use reduction operator 6 to calculate the deformation result 3 to obtain reduction result 6. Use matrix multiplication operator 23 to calculate the deformation result 3 to obtain reverse matrix multiplication result 8. Use matrix multiplication operator 24 to calculate the deformation result 3 to obtain reverse matrix multiplication result 9.

[0102] The reverse matrix multiplication result 5 output by the matrix multiplication operator 18 and the reverse matrix multiplication result 7 output by the matrix multiplication operator 21 are input to the point-to-point operator 15 for calculation to obtain the point-to-point calculation result 7.

[0103] The point-to-point calculation result 7 and the reverse matrix multiplication result 9 output by the matrix multiplication operator 24 are input to the point-to-point operator 16 for calculation to obtain the point-to-point calculation result 8.

[0104] The point-to-point calculation result 8 and the layer normalization result 4 output by the reverse layer normalization operator 2 are input to the point-to-point operator 17 for calculation to obtain the backpropagation result.

[0105] It is set that 5 templates are preset, namely template 1, template 2, template 3, template 4, and template 5. Among them, template 1 is a fused operator composed of a matrix multiplication operator, a point-to-point operator, a deformation operator, and a reordering operator; template 2 is a fused operator composed of a matrix multiplication operator, a point-to-point operator, and a point-to-point operator; template 3 is a fused operator composed of a matrix multiplication operator, a reordering operator, and a deformation operator; template 4 is a fused operator composed of a matrix multiplication operator and a point-to-point operator; template 5 is a fused operator composed of a matrix multiplication operator, a deformation operator, and a reordering operator.

[0106] Adopting the method of template matching, Figure 5 the computational graph is sliced into 26 computational subgraphs (i.e., reverse subgraphs), see Figure 6 , which are respectively: reverse sub Figure 1 to reverse subgraph 26.

[0107] In specific implementation, since the reverse layer normalization operator 1 in the computational graph does not match any of the above 5 templates, the reverse layer normalization operator 1 is separately used as the reverse sub Figure 1 . Similarly, the reduction operator 1 is used as the reverse sub Figure 2 , and the matrix multiplication operator 9 is used as the reverse sub Figure 3 .

[0108] Since the matrix multiplication operator 10 and the point-to-point operator 12 in the computational graph match template 4, the matrix multiplication operator 10 and the point-to-point operator 12 in the computational graph are divided into the reverse sub Figure 4 .

[0109] And so on, other reverse subgraphs can be obtained in the same way. The specific content of each of the other reverse subgraphs is as follows:

[0110] Reverse sub Figure 5 includes: reduction operator 2;

[0111] Reverse sub Figure 6 includes: matrix multiplication operator 11;

[0112] Reverse sub Figure 7It includes: matrix multiplication operator 12 and point-to-point operator 13;

[0113] Reverse sub- Figure 8 It includes: reverse layer normalization operator 2;

[0114] Reverse sub- Figure 9 It includes: reduction operator 3;

[0115] Reverse sub- Figure 10 It includes: matrix multiplication operator 13;

[0116] Reverse sub- Figure 11 It includes: matrix multiplication operator 14, deformation operator 5, reordering operator 5;

[0117] Reverse sub- Figure 12 It includes: matrix multiplication operator 15;

[0118] Reverse sub-graph 13 includes: reverse normalization operator;

[0119] Reverse sub-graph 14 includes: point-to-point operator 14;

[0120] Reverse sub-graph 15 includes: matrix multiplication operator 16, reordering operator 6, deformation operator 6;

[0121] Reverse sub-graph 16 includes: reduction operator 4;

[0122] Reverse sub-graph 17 includes: matrix multiplication operator 17;

[0123] Reverse sub-graph 18 includes: matrix multiplication operator 18.

[0124] Reverse sub-graph 19 includes: matrix multiplication operator 19, reordering operator 7, deformation operator 7;

[0125] Reverse sub-graph 20 includes: reduction operator 5;

[0126] Reverse sub-graph 21 includes: matrix multiplication operator 20;

[0127] Reverse sub-graph 22 includes: matrix multiplication operator 21 and point-to-point operator 15.

[0128] Reverse sub-graph 23 includes: matrix multiplication operator 22, reordering operator 8, deformation operator 8;

[0129] Reverse sub-graph 24 includes: reduction operator 6;

[0130] Reverse sub-graph 25 includes: matrix multiplication operator 23;

[0131] Reverse sub-graph 26 includes: matrix multiplication operator 24, point-to-point operator 16, point-to-point operator 17.

[0132] Consisting of Figure 5 andFigure 6 It can be seen that Figure 5 the shown computational graph contains 39 operators. After splitting the computational graph of Figure 6 , 26 computational sub-graphs are obtained, that is, the number of computational sub-graphs contained in the computational graph is less than the number of operators contained in the computational graph.

[0133] In some embodiments, for each computational sub-graph, if the computational sub-graph includes at least one operator combination for layout adjustment, the following sub-graph optimization operations are performed on the computational sub-graph:

[0134] Replace at least one operator combination for layout adjustment contained in the computational sub-graph with the output address corresponding to each operator combination for layout adjustment.

[0135] Specifically, the operator combination for layout adjustment includes: a deformation operator and a reordering operator. The deformation operator is executed before the reordering operator, or the deformation operator is executed after the reordering operator. The present application does not make specific limitations on the execution order of the deformation operator and the reordering operator in the operator combination.

[0136] Since the deformation operator and the reordering operator are IO-intensive operators, their essence is to convert one data arrangement of a tensor in memory to another data arrangement in memory. Before executing the operator combination of the deformation operator and the reordering operator, the original data arrangement (i.e., the original address in memory) of the input tensor can be obtained in advance. Therefore, the output address corresponding to the target data arrangement obtained after processing the data arrangement of the input tensor using the above operator combination can be calculated in advance. Then, the output address is merged into the output of the operator before this operator combination in the computational sub-graph, so that this operator combination can be eliminated in the computational sub-graph. After obtaining the calculation result of the previous operator, the calculation result of the previous operator can be directly output according to the output address corresponding to the target data arrangement.

[0137] For example, referring to Figure 7 , assume that in the result of computational graph splitting, the forward sub-graph includes: matrix multiplication operator 5, deformation operator 4, and reordering operator 4. The following optimization operations are performed on the forward sub-graph:

[0138] First, obtain the original data arrangement (i.e., the original address) in memory of the matrix multiplication result of matrix multiplication operator 5; then calculate the output address corresponding to the target data arrangement obtained by sequentially performing the deformation operation and the reordering operation on the original data arrangement of the matrix multiplication result by deformation operator 4 and reordering operator 4. Merge this output address into the kernel function of matrix multiplication operator 5 to eliminate deformation operator 4 and reordering operator 4 from the forward sub-graph. During the subsequent execution of the forward sub-graph, after executing matrix multiplication operator 5 to obtain the matrix multiplication result, directly output the matrix multiplication result according to the output address corresponding to the target data arrangement.

[0139] For example, referring to Figure 8 , assuming that in the computational graph partitioning result, the reverse subgraph includes: matrix multiplication operator 16, reordering operator 6, and deformation operator 6. The following optimization operations are performed on the reverse subgraph:

[0140] First, obtain the original data layout (i.e., the original address) of the matrix multiplication result of matrix multiplication operator 16 in memory; then calculate the reordering operation and deformation operation performed by reordering operator 6 and deformation operator 6 on the original data layout of the matrix multiplication result in sequence to obtain the output address corresponding to the target data layout. Merge this output address into the kernel function of matrix multiplication operator 16 to eliminate reordering operator 6 and deformation operator 6 from the reverse subgraph. During the subsequent execution of the reverse subgraph, after obtaining the matrix multiplication result by executing matrix multiplication operator 16, directly output the matrix multiplication result according to the output address corresponding to the target data layout.

[0141] In the embodiments of the present application, by replacing the operator combination for layout adjustment in the computational subgraph with the output address corresponding to the operator combination for layout adjustment, the latency caused by the deformation operator and the reordering operator reading and writing the video memory is reduced, thereby improving the computing performance.

[0142] Step 202: Obtain the subgraph merging granularity based on the resource amount of the index resources on the artificial intelligence chip.

[0143] Specifically, the index resources are used to index the video memory object for the input tensor or output tensor. When the index resource is the Usharp id for indexing the video memory object, the resource amount of the index resource refers to the number of Usharp ids. For an artificial intelligence chip, the resource amount of the index resource is limited; since the role of the index resource is to index the video memory object for the input tensor or output tensor, the total number of input tensors and output tensors included in each operator is also limited.

[0144] Based on this, the present application determines an upper limit value of the total number of input tensors and output tensors included in a merged subgraph based on the resource amount of the index resources, and uses this upper limit value as the subgraph merging granularity.

[0145] For example, if the number of Usharp ids on the artificial intelligence chip is set to 100, then the upper limit value of the total number of input tensors and output tensors included in a merged subgraph is 100.

[0146] It should be noted that in this application, the sub-graph optimization operation of the computing sub-graph can be performed first, and then the operation of determining the sub-graph merging granularity can be performed; or the operation of determining the sub-graph merging granularity can be performed first, and then the sub-graph optimization operation of the computing sub-graph can be performed; or the sub-graph optimization operation of the computing sub-graph and the operation of determining the sub-graph merging granularity can be performed in parallel; this application does not specifically limit the execution order of the sub-graph optimization operation of the computing sub-graph and the operation of determining the sub-graph merging granularity.

[0147] Step 203: Assemble and merge multiple computing sub-graphs according to the sub-graph merging granularity to obtain at least one merged sub-graph.

[0148] In some embodiments, multiple computing sub-graphs are sequentially assembled and merged according to the sub-graph merging granularity and the execution order of the multiple computing sub-graphs to obtain at least one merged sub-graph.

[0149] Specifically, traverse multiple computing sub-graphs in sequence according to the execution order of the multiple computing sub-graphs. In the initial stage, use the first traversed computing sub-graph as the reference sub-graph. Traverse the other computing sub-graphs in sequence according to the execution order of the other computing sub-graphs. When each computing sub-graph is traversed, merge this computing sub-graph with the reference sub-graph to obtain a merge result.

[0150] Judge whether the total number of input tensors and output tensors included in the merge result is less than or equal to the sub-graph merging granularity; if so, use the merge result as the new reference sub-graph and traverse the next computing sub-graph;

[0151] Otherwise, output the reference sub-graph in the merge result as a merged sub-graph; then use the computing sub-graph in the merge result as the new reference sub-graph and traverse the next computing sub-graph;

[0152] And so on, until the last computing sub-graph is traversed and stopped, and output the reference sub-graph obtained after traversing the last computing sub-graph as a merged sub-graph.

[0153] In practical applications, when merging the reference sub-graph and the computing sub-graph, there may be overlapping tensors between the output tensors of the reference sub-graph and the input tensors of the computing sub-graph. These overlapping tensors will be converted into intermediate tensors of the merge result, rather than the input tensors or output tensors of the merge result. Therefore, when counting the total number of input tensors and output tensors included in the merge result, the total number of input tensors and output tensors included in the computing sub-graph and the total number of input tensors and output tensors included in the reference sub-graph can be summed to obtain a preliminary summation result. Then, the number of tensors converted into intermediate tensors is removed from the preliminary summation result to obtain the total number of input tensors and output tensors included in the merge result. That is to say, the total number of input tensors and output tensors included in the merge result is less than or equal to the sum of the total number of input tensors and output tensors included in the computing sub-graph and the reference sub-graph respectively.

[0154] For example, set the sub-graph merging granularity on the artificial intelligence chip to 9 Usharp ids, and divide the computational graph into 4 computational sub-graphs, namely computational sub- Figure 1 to computational sub- Figure 4 , where computational sub- Figure 1 includes 3 input tensors and 3 output tensors; computational sub- Figure 2 includes 1 input tensor and 1 output tensor; computational sub- Figure 3 includes: 6 input tensors and 3 output tensors; computational sub- Figure 4 includes: 3 input tensors and 3 output tensors.

[0155] Take computational sub- Figure 1 as the reference sub-graph and merge it with computational sub- Figure 2 to obtain Merge Result 1. The total number of input tensors and output tensors included in computational sub- Figure 1 is 6, and the total number of input tensors and output tensors included in computational sub- Figure 2 is 2. The preliminary summation result obtained by adding the two is 8. Since 1 output tensor of the reference sub-graph is the same as 1 input tensor of computational sub- Figure 2 , that is, these two tensors are the intermediate tensors of Merge Result 1. Therefore, the number of intermediate tensors is removed from the preliminary summation result to obtain that the total number of input tensors and output tensors included in Merge Result 1 is 6 (less than the sub-graph merging granularity). Further, update the reference sub-graph to Merge Result 1.

[0156] Merge the reference sub-graph with computational sub- Figure 3 to obtain Merge Result 2. Since 2 output tensors of the reference sub-graph are the same as 2 input tensors of computational sub- Figure 3 , that is, these 4 tensors are converted into the intermediate tensors of Merge Result 2. Using the same method as above, it can be obtained that the total number of input tensors and output tensors included in Merge Result 2 is 11 (greater than the sub-graph merging granularity). Therefore, take the reference sub-graph in Merge Result 2 (i.e., Merge Result 1) as the merged sub- Figure 1 output.

[0157] Take computational sub- Figure 3 as the new reference sub-graph and merge it with computational sub- Figure 4 to obtain Merge Result 3. Since 3 output tensors of the reference sub-graph are the same as 3 input tensors of computational sub- Figure 4 , that is, these 6 tensors are converted into the intermediate tensors of Merge Result 3. Using the same method as above, it can be obtained that the total number of input tensors and output tensors included in Merge Result 3 is 9 (equal to the sub-graph merging granularity). Further, update the reference sub-graph to Merge Result 3.

[0158] Since computational sub-Figure 4 This is the last computational sub-graph. Therefore, the current reference sub-graph (i.e., the merged result 3) is used as the merged sub- Figure 2 output.

[0159] In the embodiments of the present application, multiple computational sub-graphs are assembled and merged according to the sub-graph merging granularity to obtain at least one merged sub-graph. In this way, the computational graph can be sliced into as few merged sub-graphs as possible. Therefore, when each merged sub-graph is compiled into a corresponding merged kernel function and executed, the number of merged kernel functions is greatly reduced, thereby reducing the time overhead of launching kernel functions.

[0160] In some embodiments, for each merged sub-graph, the following operations are respectively performed:

[0161] Traverse each intermediate tensor of a merged sub-graph. For each intermediate tensor traversed, cache resources in the artificial intelligence chip are allocated to the intermediate tensor, where the intermediate tensor is a tensor transmitted between computational sub-graphs in a merged sub-graph.

[0162] Specifically, the cache resources allocated to the intermediate tensor can be the GMB in the artificial intelligence chip.

[0163] In order to cache the entire intermediate tensor in the cache resources as much as possible to avoid fragmentation, the present application preferentially allocates cache resources to intermediate tensors that occupy less space.

[0164] Specifically, for each intermediate tensor traversed and when it is determined that the storage space occupied by the intermediate tensor is less than a preset threshold, cache resources in the artificial intelligence chip are allocated to the intermediate tensor.

[0165] The preset threshold can be equal to the size of the cache resources on the artificial intelligence chip or less than the size of the cache resources on the artificial intelligence chip.

[0166] In some embodiments, in order to effectively reuse the cache resources in the artificial intelligence chip on the premise of ensuring the correct calculation of the operator and improve the utilization rate of the cache resources. The present application allocates cache resources to intermediate tensors based on the life cycle of each intermediate tensor.

[0167] When the life cycle of the intermediate tensor ends, the cache resources allocated to the intermediate tensor are recycled and then reallocated to other intermediate tensors to achieve cache resource reuse.

[0168] When an intermediate tensor depends on multiple computational subgraphs, the lifecycle of the intermediate tensor ends only after all the computational subgraphs it depends on have been executed. If there are other intermediate tensors between the computational subgraphs that the intermediate tensor depends on, to avoid data loss and ensure the correctness of the computation, the cache resources allocated for this intermediate tensor do not overlap with the cache resources allocated for other intermediate tensors between the computational subgraphs that this intermediate tensor depends on. When the lifecycle of this intermediate tensor ends, the cache resources allocated to this intermediate tensor are reused.

[0169] For example, refer to Figure 9 , the merged subgraph includes computational sub- Figure 1 , computational sub- Figure 2 and computational sub- Figure 3 . The intermediate tensors of the merged subgraph include: the intermediate tensor F1 between computational sub- Figure 1 and computational sub- Figure 2 , the intermediate tensor F1 between computational sub- Figure 1 and computational sub- Figure 3 , and the intermediate tensor F2 between computational sub- Figure 2 and computational sub- Figure 3 .

[0170] For the intermediate tensor F1 between computational sub- Figure 1 and computational sub- Figure 2 , the first storage space in the GMB is allocated. Since the intermediate tensor F1 needs to be input not only to computational sub- Figure 2 for computation, but also to computational sub- Figure 3 for computation, that is, the intermediate tensor F1 depends not only on computational sub- Figure 2 , but also on computational sub- Figure 3 , the lifecycle of the intermediate tensor F1 ends after computational sub- Figure 3 is executed. Therefore, the first storage space allocated for the intermediate tensor F1 cannot be allocated to the intermediate tensor F2 between the computational sub- Figure 2 that the intermediate tensor F1 depends on and computational sub- Figure 3 . Instead, the second storage space in the GMB needs to be allocated for the intermediate tensor F2, and the first storage space and the second storage space are different storage spaces in the GMB.

[0171] For example, refer to Figure 10 , the merged subgraph includes computational sub- Figure 1 , computational sub- Figure 2 and computational sub- Figure 3 . The intermediate tensors of the merged subgraph include: the intermediate tensor F1 between computational sub- Figure 1 and computational sub- Figure 2 , and the intermediate tensor F2 between computational sub- Figure 2 and computational sub- Figure 3 .

[0172] For the computing unit Figure 1 and the computing unit Figure 2 allocate the first storage space in the GMB for the intermediate tensor F1 between them. Since the lifecycle of the intermediate tensor F1 ends when the computing unit Figure 2 finishes execution, therefore, the first storage space in the GMB can be reused to save the intermediate tensor F2.

[0173] Step 204: Compile each merged sub-graph into a corresponding merged kernel function, and execute the merged kernel function corresponding to each merged sub-graph to obtain the calculation result of the encoding module.

[0174] Specifically, the central processing unit in the computer device compiles each merged sub-graph into a merged kernel function (MegaKernel), and emits the merged kernel function to the artificial intelligence chip in the computer device. The artificial intelligence chip sequentially executes the corresponding merged kernel functions according to the execution order of each merged sub-graph to obtain the calculation result of the encoding module. In addition, the process of obtaining the merged sub-graphs described in the above steps 201 to 203 is also executed by the central processing unit in the computer device.

[0175] In the embodiments of the present application, the computational graph of the encoding module is sliced into multiple computing sub-graphs based on a preset template; then, based on the resource amount of the index resources on the artificial intelligence chip, the sub-graph merging granularity is obtained; and then, according to the sub-graph merging granularity, the multiple computing sub-graphs are assembled and merged to obtain at least one merged sub-graph, that is, considering the index resources on the artificial intelligence chip, the computational graph is sliced into as few merged sub-graphs as possible. Therefore, when each merged sub-graph is compiled into a corresponding merged kernel function and executed, the number of merged kernel functions is greatly reduced, thereby reducing the time overhead of emitting kernel functions. Secondly, the computational graph optimization method of the present application not only considers the operators included in the computational graph itself, but also combines the resource situation of the underlying chip that actually executes the computational graph, so as to maximize the optimization of the computational graph and improve the optimization effect of the computational graph.

[0176] Based on the same technical concept, the embodiments of the present application provide a structural schematic diagram of a computational graph optimization device, as Figure 11 shown. The computational graph optimization device 1100 includes:

[0177] A slicing module 1101, configured to slice the computational graph of the encoding module into multiple computing sub-graphs based on a preset template;

[0178] A merging module 1102, configured to obtain the sub-graph merging granularity based on the resource amount of the index resources on the artificial intelligence chip; and assemble and merge the multiple computing sub-graphs according to the sub-graph merging granularity to obtain at least one merged sub-graph;

[0179] An execution module 1103, configured to compile each merged sub-graph into a corresponding merged kernel function, and execute the merged kernel function corresponding to each said merged sub-graph, to obtain a calculation result of the encoding module.

[0180] Optionally, the index resource is used to: index a video memory object for an input tensor or an output tensor;

[0181] The merging module 1102 is specifically configured to:

[0182] Based on the resource amount of the index resource, determine an upper limit value of the total number of input tensors and output tensors included in a merged sub-graph;

[0183] Use the upper limit value as the sub-graph merging granularity.

[0184] Optionally, the merging module 1102 is specifically configured to:

[0185] Assemble and merge the multiple computational sub-graphs in sequence according to the sub-graph merging granularity and the execution order of the multiple computational sub-graphs, to obtain at least one merged sub-graph.

[0186] Optionally, the merging module 1102 is further configured to:

[0187] After assembling and merging the multiple computational sub-graphs according to the sub-graph merging granularity to obtain at least one merged sub-graph, for each merged sub-graph, respectively perform the following operations:

[0188] Traverse each intermediate tensor of a merged sub-graph, and for each traversed intermediate tensor, allocate cache resources in the artificial intelligence chip for the one intermediate tensor, where the intermediate tensor is a tensor passed between computational sub-graphs in the one merged sub-graph.

[0189] Optionally, the merging module 1102 is specifically configured to:

[0190] For each traversed intermediate tensor, and when it is determined that the storage space occupied by the one intermediate tensor is less than a preset threshold, allocate the cache resources for the one intermediate tensor.

[0191] Optionally, the cache resources allocated for the one intermediate tensor do not overlap with the cache resources allocated for other intermediate tensors between computational sub-graphs on which the one intermediate tensor depends.

[0192] Optionally, the merging module 1102 is further configured to:

[0193] Before obtaining at least one merged sub-graph by assembling and merging the multiple computing sub-graphs according to the sub-graph merging granularity, if the computing sub-graph includes at least one operator combination for layout adjustment, the following sub-graph optimization operations are performed on the computing sub-graph:

[0194] Replace the at least one operator combination for layout adjustment included in the computing sub-graph with the respective output addresses corresponding to the at least one operator combination for layout adjustment.

[0195] In the embodiments of the present application, the computing graph of the encoding module is segmented into multiple computing sub-graphs based on a preset template; then, the sub-graph merging granularity is obtained based on the resource amount of the index resources on the artificial intelligence chip; and then, at least one merged sub-graph is obtained by assembling and merging the multiple computing sub-graphs according to the sub-graph merging granularity. That is, considering the index resources on the artificial intelligence chip, the computing graph is segmented into as few merged sub-graphs as possible. Therefore, when each merged sub-graph is compiled into a corresponding merged kernel function and executed, the number of merged kernel functions is greatly reduced, thereby reducing the time overhead of launching kernel functions. Secondly, the computing graph optimization method of the present application not only considers the operators included in the computing graph itself, but also combines the resource conditions of the underlying chip that actually executes the computing graph, so as to maximize the optimization of the computing graph and improve the optimization effect of the computing graph.

[0196] Based on the same technical concept, the embodiments of the present application provide a computer device, such as Figure 12 shown, including at least one artificial intelligence chip 100 and a memory 1201 connected to the at least one artificial intelligence chip 100. In the embodiments of the present application, the specific connection medium between the artificial intelligence chip 100 and the memory 1201 is not limited. Figure 12 Taking the example that the artificial intelligence chip 100 and the memory 1201 are connected by a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0197] In the embodiments of the present application, the memory 1201 stores instructions executable by the at least one artificial intelligence chip 100. The at least one artificial intelligence chip 100 can execute the steps of the above computing graph optimization method by executing the instructions stored in the memory 1201.

[0198] Among them, the artificial intelligence chip 100 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and circuits, and achieve computing graph optimization by running or executing instructions stored in the memory 1201 and calling data stored in the memory 1201. Optionally, the artificial intelligence chip 100 may include one or more processing units. The artificial intelligence chip 100 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the artificial intelligence chip 100. In some embodiments, the artificial intelligence chip 100 and the memory 1201 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.

[0199] The artificial intelligence chip 100 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0200] The memory 1201, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1201 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 1201 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited to this. The memory 1201 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, used to store program instructions and / or data.

[0201] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium that stores a computer program executable by a computer device. When the computer program runs on the computer device, it causes the computer device to execute the steps of the above-mentioned computing graph optimization method.

[0202] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, they cause the computer device to execute the steps of the above-mentioned computing graph optimization method.

[0203] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0204] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer device or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0206] These computer program instructions can also be loaded onto a computer device or other programmable data processing device, such that a series of operational steps are executed on the computer device or other programmable device to produce a process implemented by the computer device, so that the instructions executed on the computer device or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0207] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0208] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A computational graph optimization method, characterized in that, Including: Segmenting the computational graph of an encoding module into multiple computational subgraphs based on a preset template; Determining an upper limit value of the total number of input tensors and output tensors included in a merged subgraph based on the resource amount of index resources on an artificial intelligence chip, where the index resources are used for indexing video memory objects for input tensors or output tensors; Using the upper limit value as the subgraph merging granularity; Assembling and merging the multiple computational subgraphs according to the subgraph merging granularity to obtain at least one merged subgraph; Compiling each merged subgraph into a corresponding merged kernel function and executing the merged kernel function corresponding to each merged subgraph to obtain the computational result of the encoding module.

2. The method according to claim 1, wherein The step of assembling and merging the multiple computational subgraphs according to the subgraph merging granularity to obtain at least one merged subgraph includes: Assembling and merging the multiple computational subgraphs in sequence according to the subgraph merging granularity and the execution order of the multiple computational subgraphs to obtain at least one merged subgraph.

3. The method according to claim 1, characterized in that After the step of assembling and merging the multiple computational subgraphs according to the subgraph merging granularity to obtain at least one merged subgraph, it further includes: For each merged subgraph, respectively perform the following operations: Traversing each intermediate tensor of a merged subgraph, and for each traversed intermediate tensor, allocating cache resources in the artificial intelligence chip to the intermediate tensor, where the intermediate tensor is a tensor transmitted between computational subgraphs in the merged subgraph.

4. The method according to claim 3, characterized in that, The step of for each traversed intermediate tensor, allocating cache resources in the artificial intelligence chip to the intermediate tensor includes: For each traversed intermediate tensor, when it is determined that the storage space occupied by the intermediate tensor is less than a preset threshold, allocating the cache resources to the intermediate tensor.

5. The method according to claim 3, wherein It further includes: The cache resources allocated to an intermediate tensor do not overlap with the cache resources allocated to other intermediate tensors between the computational subgraphs on which the intermediate tensor depends.

6. The method according to any one of claims 1 to 5, characterized in that Before the step of assembling and merging the multiple computational subgraphs according to the subgraph merging granularity to obtain at least one merged subgraph, it further includes: For each computational subgraph, if the computational subgraph includes at least one operator combination for layout adjustment, perform the following subgraph optimization operation on the computational subgraph: Replacing the at least one operator combination for layout adjustment included in the computational subgraph with the output addresses corresponding to the at least one operator combination for layout adjustment respectively.

7. A computer device, comprising a memory, an artificial intelligence chip, and a computer program stored on the memory and executable on the artificial intelligence chip, characterized in that, When the artificial intelligence chip executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 6.

9. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Segmentation method, device and equipment of neural network calculation graph and storage medium

    CN117576125A