Tensor program optimization method and device

By performing computational graph analysis, transformation and kernel splitting of the tensor program of the deep neural network model, the problem of increasing processing time of large tensors is solved, and more efficient execution efficiency and memory utilization are achieved.

CN120234513APending Publication Date: 2025-07-01TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224540.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

As the size of deep neural network models grows, the size of intermediate tensors changes significantly, resulting in increased processing time, and the prior art lacks effective solutions to reduce the processing time of tensors.

Method used

Tensor attribute analysis is obtained by a tensor attribute based on the calculation graph and preset tensor attributes of the tensor program, and the tensor attributes of each tensor in the calculation graph are obtained; tensor transformation is performed based on these properties and preset transformation rules to obtain an equivalent graph; then the calculation kernel is split on the equivalent graph to obtain multiple calculation kernels.

Benefits of technology

This method can improve the execution efficiency of tensor programs, reduce memory consumption, and optimize the computing graph structure to reduce memory access overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234513A_ABST
    Figure CN120234513A_ABST
Patent Text Reader

Abstract

The invention provides a tensor program optimization method and device, and the method comprises the steps: carrying out the tensor attribute analysis based on a calculation graph of a tensor program and a preset tensor attribute, and obtaining the tensor attribute of each tensor in the calculation graph; performing tensor transformation based on the tensor attribute of each tensor in the computational graph and a transformation rule to obtain an equivalent graph of the computational graph; wherein the transformation rule is preset; and performing calculation kernel splitting on the equivalent graph of the calculation graph to obtain a plurality of calculation kernels corresponding to the calculation graph. The device is used for executing the method. According to the tensor program optimization method and device provided by the embodiment of the invention, the execution efficiency of the tensor program is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to an optimization method and device for a tensor program. Background Art

[0002] Currently, Deep Neural Network (DNN) has played an important role in fields such as healthcare, autonomous driving, and natural language processing. Learning from large-scale datasets and solving complex problems through DNN demonstrates its wide adaptability and effectiveness.

[0003] The scale of DNN models shows an obvious growth trend. The growth of the model scale is divided into two different categories: (1) Parameter Determined Axis (P-axis). The P-axis mainly designs the size of the hidden layer of DNN, and increasing the size of the hidden layer can improve the quality of the inference result. (2) Input Determined Axis (I-axis). The I-axis refers to dimensions such as sequence length and image size. Increasing these dimensions enables DNN to process larger-scale input information. The growth rates of the P-axis and I-axis are different, and the difference in the growth rates of different axes leads to a significant change in the size of the intermediate tensor. As the difference between the sizes of the P-axis and I-axis becomes larger, the generated tensor also becomes larger, resulting in an increase in processing time. For the processing of the enlarged tensor, there is currently no effective solution to reduce the processing time of the tensor. Summary of the Invention

[0004] In view of the problems in the prior art, embodiments of the present invention provide an optimization method and device for a tensor program.

[0005] In a first aspect, the present invention proposes an optimization method for a tensor program, including:

[0006] Performing tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph;

[0007] Performing tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; wherein, the transformation rules are preset;

[0008] Performing computational kernel splitting on the equivalent graph of the computational graph to obtain a plurality of computational kernels corresponding to the computational graph.

[0009] In a second aspect, the present invention provides an optimization device for a tensor program, including:

[0010] An analysis module, configured to perform tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph;

[0011] A transformation module, configured to perform tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; wherein the transformation rules are preset.

[0012] A splitting module, configured to split the equivalent graph of the computational graph into computational kernels to obtain a plurality of computational kernels corresponding to the computational graph.

[0013] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the program to implement the optimization method for a tensor program according to any one of the above embodiments.

[0014] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program / instructions, and when the computer program / instructions are executed by a processor, the optimization method for a tensor program according to any one of the above embodiments is implemented.

[0015] In a fifth aspect, the present invention provides a computer program product including a computer program / instructions, and when the computer program / instructions are executed by a processor, the optimization method for a tensor program according to any one of the above embodiments is implemented.

[0016] The optimization method and device for a tensor program provided by the embodiments of the present invention can perform tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; perform tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; split the equivalent graph of the computational graph into computational kernels to obtain a plurality of computational kernels corresponding to the computational graph, which can improve the execution efficiency of the tensor program. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0018] Figure 1 is a simplified attention computational graph provided by the first embodiment of the present invention.

[0019] Figure 2 is a schematic flowchart of the optimization method for a tensor program provided by the second embodiment of the present invention.

[0020] Figure 3 is a schematic diagram of obtaining tensor attributes provided by the third embodiment of the present invention.

[0021] Figure 4 It is a schematic flowchart of an optimization method for a tensor program provided by the fourth embodiment of the present invention.

[0022] Figure 5 It is a schematic flowchart of an optimization method for a tensor program provided by the fifth embodiment of the present invention.

[0023] Figure 6 It is a schematic flowchart of an optimization method for a tensor program provided by the sixth embodiment of the present invention.

[0024] Figure 7 It is a schematic flowchart of an optimization method for a tensor program provided by the seventh embodiment of the present invention.

[0025] Figure 8 It is an execution effect diagram of a tensor program provided by the eighth embodiment of the present invention.

[0026] Figure 9 It is a schematic structural diagram of an optimization device for a tensor program provided by the ninth embodiment of the present invention.

[0027] Figure 10 It is a schematic structural diagram of an optimization device for a tensor program provided by the tenth embodiment of the present invention.

[0028] Figure 11 It is a schematic structural diagram of an optimization device for a tensor program provided by the eleventh embodiment of the present invention.

[0029] Figure 12 It is a schematic structural diagram of an optimization device for a tensor program provided by the twelfth embodiment of the present invention.

[0030] Figure 13 It is a schematic structural diagram of an optimization device for a tensor program provided by the twelfth embodiment of the present invention.

[0031] Figure 14 It is a schematic physical structure diagram of a computer device provided by the fourteenth embodiment of the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention. It should be noted that, without conflict, the embodiments and features in the embodiments in this application can be combined with each other arbitrarily.

[0033] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant countries and regions, necessary confidentiality measures are taken, it does not violate public order and good customs, and a corresponding operation entry is provided for users to choose to authorize or reject.

[0034] To facilitate the understanding of the technical solution provided in this application, the relevant content of the technical solution of this application will be described first below.

[0035] Language Model (LLM for short): A natural language processing model based on deep learning, whose goal is to generate natural language text that conforms to grammar and semantic rules.

[0036] Deep neural network: A neural network with a certain depth stacked by multiple neural network layers.

[0037] Tensor: A special data structure similar to data and matrices, with characteristics such as dimensions, offsets, and shapes.

[0038] Operator: A computational unit used to perform various mathematical operations and operations.

[0039] Fusion: Combining multiple computational units into one computational unit to complete the calculation, reducing the operation of reading and writing intermediate data to memory, thereby saving calculation time.

[0040] Computation Graph: A graph structure composed of operators and tensors, where the operators are the nodes on the graph, and the input and output tensors are the edges on the graph.

[0041] Tensor Compiler: A compiler for neural networks that generates higher inference performance through the optimization of the computation graph and operator optimization.

[0042] Computation Graph Optimization: Refers to the optimization technology on the computation graph, which transforms the graph structure into a computation graph with the same computation semantics but better computation performance.

[0043] Kernel: The basic computational unit that executes operations on the GPU, written in the form of cuda or ptx code on NVIDIA GPUs, and can be executed after being compiled by the compiler provided by the manufacturer.

[0044] Dataflow Analysis: A technique for collecting information about the values computed by a program at different points, obtaining static information about the entire program through forward or backward propagation.

[0045] The growth rates of the parameter-determined axis (P-axis) and the input-determined axis (I-axis) are different. For example, from Llama-65B to Llama-3.1-405B, the P-axis (hidden layer size) increases from 8k to 16k, tripling; in contrast, the I-axis (context length) expands from 2k to 128k, increasing by 64 times. The difference in the growth rates of different axes leads to a significant change in the size of intermediate tensors.

[0046] As Figure 1 shown, a simplified attention computation graph is presented, which includes a MatMul operation followed by two Reduce operations in different directions: Reduce 0 and Reduce 1. When the sizes of the P-axis and the I-axis are similar, the input and output shapes of the MatMul operation are similar. However, as the difference in the sizes of the P-axis and the I-axis becomes larger, the output of the MatMul operation becomes very large compared to other tensors in the computation graph. These very large tensors require a large amount of processing time, making their efficiency crucial for overall performance. However, due to the operations associated with them being potentially very complex, the prior art fails to provide an efficient solution for these extremely large tensors.

[0047] As Figure 1 shown, the output tensor of MatMul is extremely large, and one of the effective ways to eliminate this tensor and reduce memory overhead is through fusion operations. However, this tensor is subsequently used by two Reduce operations, one performing reduction on each row and the other on each column. If a fusion operation is applied, the complex dependencies caused by the different reduction dimensions of the two Reduce operations will lead to low parallelism, seriously affecting performance.

[0048] This application seeks an efficient computation graph optimization scheme by analyzing the fine-grained attributes of intermediate tensors in the computation graph (such as reduction dependencies on certain axes, etc.). Although coarse-grained tensor attributes like tensor size have been considered in existing work, existing deep learning compilers ignore finer-grained tensor attributes and mainly focus on attributes related to operator operators, such as the mathematical properties of the operator operator itself or the computing resources required for the operation. Additionally, manually developed operator libraries are also unable to optimize new model variants. To address these issues, based on fine-grained tensor attributes, this application proposes an optimization method for tensor programs that can optimize large tensors, reduce memory consumption, and improve the execution efficiency of tensor programs.

[0049] Figure 2It is a schematic flowchart of the optimization method for a tensor program provided by the second embodiment of the present invention. As Figure 2 shown, the optimization method for a tensor program provided by an embodiment of the present invention includes:

[0050] S201. Perform tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph;

[0051] Specifically, performing tensor attribute analysis on the computational graph of the tensor program according to the preset tensor attributes can obtain the tensor attributes of each tensor in the computational graph. The preset tensor attributes are tensor attributes that are preset and play an important role in the optimization of the tensor program, and are set according to actual needs, which are not limited in the embodiments of the present invention.

[0052] For example, in order to obtain the tensor attributes of all tensors in the computational graph of the tensor program, a tensor attribute analysis strategy based on data flow analysis can be used to obtain the tensor attributes of each tensor.

[0053] For example, as shown in Table 1, the preset tensor attributes include dimension attributes and tensor overall attributes. The dimension attributes include reduction dependency and broadcast. The reduction dependency corresponds to the reduction dependency of the tensor dimension, which helps to analyze the parallel efficiency on this dimension; the broadcast shows whether this dimension will be subject to a broadcast operation by a certain operator. The reduction dependency includes non-parameter (NonPara), batch, and reuse. Non-parameter represents a non-parallelizable dimension, batch represents an easily parallelizable dimension, and reuse represents a parallelizable dimension with data reuse. Among them, the priority of non-parameter is higher than that of batch, and the priority of batch is higher than that of reuse.

[0054] The tensor overall attributes include size and value. The size is an integer representing the total size of the tensor. The value is divided into seven categories: Var, Zero, PosConst, NegConst, PosInf, NegInf, and NaN. Var indicates that the tensor value is undetermined, Zero indicates that the tensor is all zero, PosConst indicates that the tensor is a positive constant, NegConst indicates that the tensor is a negative constant, PosInf indicates that the tensor is positive infinity, NegInf indicates that the tensor is negative infinity, and NaN indicates that the tensor is not a number.

[0055] Table 1 Preset Tensor Attributes

[0056]

[0057]

[0058] S202. Perform tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; wherein, the transformation rules are preset.

[0059] Specifically, after obtaining the tensor attributes of all tensors in the computational graph of the tensor program, tensor transformation can be performed according to the tensor attributes of each tensor in the computational graph and transformation rules, and the tensors in the computational graph are replaced with the transformed tensors to obtain an equivalent graph of the computational graph. The transformation of tensors mainly considers the impact of the size of intermediate tensors in the computational graph on memory access. Through the transformation rules, the size of intermediate tensors can be reduced, thereby reducing the memory access overhead and improving the overall performance. The transformation rules are preset and are set according to actual needs, which are not limited in the embodiments of the present invention.

[0060] For example, as shown in Table 2, A, B, and C represent tensors. The rules column corresponds to multiple transformation conditions, and each transformation condition has a corresponding legal condition. When the legal condition is no, it means the legal condition is empty; when the legal condition is yes, it means there is a legality judgment condition and a legality judgment needs to be made.

[0061] Table 2 Transformation Rules

[0062]

[0063]

[0064] S203. Perform computational kernel splitting on the equivalent graph of the computational graph to obtain multiple computational kernels corresponding to the computational graph.

[0065] Specifically, after obtaining the equivalent graph of the computational graph, performing computational kernel splitting on the equivalent graph of the computational graph can split the equivalent graph of the computational graph into multiple computational kernels. Each computational kernel will be handed over to the GPU to execute the corresponding calculation. The above multiple computational kernels can be a combination including non-convex kernels. Non-convex kernels have better parallelism and square efficiency, which can improve the execution efficiency of the tensor program.

[0066] The optimization method of the tensor program provided by the embodiments of the present invention can perform tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; perform tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; perform computational kernel splitting on the equivalent graph of the computational graph to obtain multiple computational kernels corresponding to the computational graph, which can improve the execution efficiency of the tensor program.

[0067] Based on the above embodiments, further, the tensor attribute analysis based on the computational graph of the tensor program and the preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph includes:

[0068] If the computational graph of the tensor program includes a single operator, then based on the computational semantics of the single operator and the preset tensor attributes, obtain the tensor attributes of the input tensors and output tensors of the single operator.

[0069] Specifically, the computational graph of the tensor program may include a single operator or multiple operators. If the computational graph of the tensor program includes a single operator, then the input tensor attributes and output tensor attributes of the single operator can be obtained according to the computational semantics of the single operator and the preset tensor attributes.

[0070] For example, the single operator included in the computational graph of the tensor program is the Batched GEMM operator. The three-dimensional input tensors of the Batched GEMM operator are A and B, and the three-dimensional output tensor is C. The expression of the Batched GEMM operator is: C[b,i,j] = ∑ k A[b,i,k] × B[b,k,j]. For the input tensors A and B, since the k dimension appears in the summation symbol, the reduction dependence of the k dimension is identified as NonPara; for the i dimension and j dimension of the tensors A, B, and C, since they do not appear in the summation symbol of the operator expression and do not exist simultaneously in all input and output tensors, the reduction dependence of the i dimension and j dimension is identified as Reuse; for the b dimension of the tensors A, B, and C, it does not appear in the summation symbol of the Batched GEMM operator expression but exists simultaneously in all input and output tensors, so the reduction dependence of this b dimension is identified as Batch. For the broadcast tensor attribute, it is only necessary to check whether the operator will broadcast a certain input tensor. If it broadcasts, the dimension of the corresponding output tensor that is broadcast is identified as Yes of the broadcast tensor attribute. For other tensor attributes, they can be directly determined by combining the shape and value of the tensor itself.

[0071] Based on the above embodiments, further, the tensor attribute analysis based on the computational graph of the tensor program and the preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph includes:

[0072] If the computational graph of the tensor program includes multiple operators, then perform iterative calculations on each operator included in the computational graph until there is no situation where a high-priority value of the same tensor attribute replaces a low-priority value in the input tensors and output tensors of all operators included in the computational graph.

[0073] Specifically, if the computation graph of a tensor program includes multiple operators, then iterate through each operator included in the computation graph on the computation graph for iterative calculation. After each iterative calculation is completed, check whether there are different values for the same tensor attributes among the tensor attributes included in the input tensors and output tensors of each operator. If there are different values for the same tensor attributes, then replace the lower-priority values with the highest-priority values of the same tensor attributes. If after a certain iterative calculation is completed, for all operators, there is no situation where the highest-priority values of the same tensor attributes are used to replace the lower-priority values, then the iteration ends. Among them, the priority relationship of the values corresponding to the tensor attributes is preset.

[0074] For example, as Figure 3 shown, two input tensors, one with a shape of N×d×h and the other with a shape of d×N′×h, are input into a matrix multiplication operator (MatMul), and a tensor with a shape of N×N′×h is output. This output tensor is fed into a reduction operator (Reduce), and finally a tensor with a shape of N×c×h is output.

[0075] The table under the matrix multiplication operator in the figure shows the reduction dependencies of the input tensors and output tensors of the matrix multiplication operator on the dimensions h, N, and N′. The reduction dependency on dimension h is "Batch"; the reduction dependencies on dimensions N and N′ are "Reuse".

[0076] The table under the reduction operator in the figure shows the reduction dependencies of the input tensors and output tensors of the reduction operator on the dimensions h, N, and N′. The reduction dependencies on dimensions h and N are "Batch". The reduction dependency on dimension N′ is "NonPara".

[0077] The table at the bottom of the figure shows how to aggregate (Aggregation) the reduction dependencies of the matrix multiplication operator and the reduction operator in each dimension. The rule is: for tensor attributes of the same dimension, take the attribute value with the highest priority. For dimension h, "Batch + Batch = Batch", for dimension N, "Reuse + Batch = Reuse", and for dimension N′, "Reuse + NonPara = NonPara". Selecting the highest-priority value can preserve the dependency relationship in that dimension, thereby avoiding parallelizing dependent dimensions or ignoring optimization opportunities for data reuse in subsequent code generation.

[0078] Figure 4 is a schematic flowchart of the optimization method for a tensor program provided in the fourth embodiment of the present invention. As Figure 4 shown, on the basis of the above embodiments, further, the obtaining of the equivalent graph of the computation graph by performing tensor transformation based on the tensor attributes of each tensor in the computation graph and transformation rules includes:

[0079] S401. Traverse each operator in the computational graph, and based on the input tensor and output tensor of each operator and the transformation rule, obtain the operators that satisfy the transformation rule and the transformed expressions corresponding to each operator.

[0080] Specifically, for each operator in the computational graph, match the transformation rule based on the input tensor and output tensor of the operator to determine whether the operator satisfies the transformation conditions in the transformation rule. If the operator does not require a legality judgment, then the expression of the operator can be transformed based on the satisfied transformation conditions. If the operator does not satisfy all the transformation conditions in the transformation rule, then the expression will not be transformed. If the operator requires a legality judgment, for the operator that satisfies the legality conditions, the expression of the operator will be transformed based on the satisfied transformation conditions; for the operator that does not satisfy the legality conditions, the expression will not be transformed. After traversing all the operators in the computational graph, the operators that satisfy the transformation rule and the transformed input tensors and output tensors corresponding to each operator can be obtained. Among them, the transformation rule is preset and is used to reduce the size of the intermediate tensors in the computational graph, which can reduce the memory access overhead and improve the overall performance.

[0081] It can be understood that during the execution of the computational graph or algorithm, starting from the input tensor, after being processed by a series of operators (such as matrix multiplication, convolution, activation function operations, etc.), all the tensors generated before obtaining the final output tensor can be called intermediate tensors.

[0082] S402. Update the computational graph based on each operator that satisfies the transformation rule and its corresponding transformed expression until a computational graph that satisfies the preset conditions is obtained as the equivalent graph of the computational graph.

[0083] Specifically, in order to prevent the equivalent graph of the computational graph from falling into a local optimal solution, that is, failing to find the computational graph with the smallest global intermediate tensor size. The computational graph can be updated based on at least one operator that satisfies the transformation rule and its corresponding transformed expression, and determine whether the updated computational graph satisfies the preset conditions. If it does not satisfy the preset conditions, continue to update the computational graph until a computational graph that satisfies the preset conditions is obtained as the equivalent graph of the computational graph.

[0084] Figure 5 It is a schematic flowchart of the optimization method for tensor programs provided in the fifth embodiment of the present invention. As Figure 5 shown, on the basis of the above embodiments, further, the obtaining the operators that satisfy the transformation rule and the transformed expressions corresponding to each operator based on the input tensor and output tensor of each operator and the transformation rule includes:

[0085] S501. Obtain the transformed expression of the operator under each matching transformation condition based on the input tensor of the operator and the transformation rule; wherein, the transformation rule includes at least one transformation condition, and each transformation condition has a corresponding legality judgment condition.

[0086] Specifically, according to the input tensor of the operator and the transformation rule, determine which transformation condition included in the transformation rule the operator matches, that is, substitute the input tensor of the operator into the expression of the transformation condition. If the expression of the transformation condition holds after substituting the input tensor, then the operator matches the transformation condition. Substitute the transformed expression of the input tensor of the operator into the expression of the operator, and the transformed expression of the operator under the matching transformation condition can be obtained. If the operator matches multiple transformation conditions, then obtain the transformed expressions of the operator under each matching transformation condition. The transformation rule includes at least one transformation condition, and each transformation condition has a corresponding legality judgment condition.

[0087] It can be understood that if the operator does not match all the change conditions, then the expression of the operator will not be changed.

[0088] For example, as shown in Table 2, the transformation rule includes 11 transformation conditions. For operator X, the input tensors of operator X are a and b. If a + b = b + a, then operator X matches the transformation condition A + B = B + A. Substitute b + a for a + b in the operator, and the transformed expression of the operator under the transformation condition A + B = B + A can be obtained.

[0089] S502. If the sum of the sizes of the transformed input tensor and output tensor of the operator under the matching transformation condition is less than the sum of the sizes of the input tensor and output tensor of the operator before transformation and the legality judgment condition corresponding to the transformation condition matched by the operator is empty, then regard the operator as an operator that meets the transformation rule.

[0090] Specifically, obtain the transformed input tensor and output tensor of the operator under the matching transformation condition, and calculate the sum S1 of the size of the transformed input tensor of the operator under the matching transformation condition and the size of the transformed output tensor of the operator under the matching transformation condition. Obtain the input tensor and output tensor of the operator before transformation, and calculate the sum S2 of the size of the input tensor of the operator before transformation and the size of the output tensor of the operator before transformation. Compare S1 and S2. If S1 is less than S2, and the legality judgment condition corresponding to the transformation condition matched by the operator is empty, then regard the operator as an operator that meets the transformation rule.

[0091] It is understandable that if S1 is greater than or equal to S2, then the operator will not be regarded as an operator that satisfies the transformation rule.

[0092] S503. If the sum of the sizes of the transformed input tensor and the output tensor of the operator under the transformation condition is smaller than the sum of the sizes of the input tensor and the output tensor of the operator before the transformation, and the legality judgment condition corresponding to the transformation condition corresponding to the operator is not empty, then based on the legality judgment condition corresponding to the transformation condition corresponding to the operator and the tensor attributes of the input tensor of the operator, determine whether the transformation of the operator under the transformation condition is legal;

[0093] Specifically, compare S1 and S2. If S1 is smaller than S2, and the legality judgment condition corresponding to the transformation condition matched by the operator is not empty, then obtain the tensor attributes of the input tensor of the operator, and based on the tensor attributes of the input tensor of the operator and the legality judgment condition corresponding to the transformation condition corresponding to the operator, determine whether the transformation of the operator under the transformation condition is legal.

[0094] S504. If the transformation of the operator under the transformation condition is legal, then regard the operator as an operator that satisfies the transformation rule.

[0095] Specifically, if the transformation of the operator under the transformation condition is legal, then regard the operator as an operator that satisfies the transformation rule. If the transformation of the operator under the transformation condition is not legal, then the operator will not be regarded as an operator that satisfies the transformation rule.

[0096] For example, as shown in Table 2, for operator Y, the input tensors of operator Y are c, d, and e. If (c / d)e = (c e) / d, then operator Y matches the transformation condition (A / B)C = (AC) / B, and the legality judgment condition corresponding to the transformation condition (A / B)C = (AC) / B is that the broadcast attribute value of B in dimension i is yes. Obtain the broadcast under the dimension attribute included in the tensor attributes of the input tensor d of operator Y. If the corresponding attribute value of the broadcast is yes, then regard operator Y as an operator that satisfies the transformation rule. If the corresponding attribute value of the broadcast is no, then the operator Y will not be regarded as an operator that satisfies the transformation rule.

[0097] Figure 6 It is a schematic flowchart of the optimization method of the tensor program provided by the sixth embodiment of the present invention. As Figure 6 shown, on the basis of the above embodiments, further, update the computational graph based on each operator that satisfies the transformation rule and their respective corresponding transformed expressions until a computational graph that satisfies the preset conditions is obtained as the equivalent graph of the computational graph, including:

[0098] S601. Obtain the transformed expressions corresponding to at least one operator that satisfies the transformation rule, update the computational graph, and obtain the current computational graph.

[0099] Specifically, the computational graph includes operators. The computational graph can be updated according to the transformed expressions corresponding to at least one operator that satisfies the transformation rule to obtain the current computational graph. The current computational graph is the computational graph that is being optimized.

[0100] S602. Calculate the difference between the tensor size of the initial computational graph and the tensor size of the current computational graph to obtain the tensor change amount; where the first initial computational graph uses the computational graph.

[0101] Specifically, obtain the tensor size K1 of the initial computational graph and the tensor size K2 of the current computational graph, and calculate the difference between K1 and K2 as the tensor change amount. The specific calculation process of the tensor size of the computational graph is prior art and will not be elaborated here. The initial computational graph will change during the process of obtaining the equivalent graph of the computational graph. The first initial computational graph uses the computational graph of the tensor program.

[0102] S603. Obtain an intermediate index value based on the tensor change amount and the maximum tensor value of the current computational graph.

[0103] Specifically, an intermediate index value can be obtained based on the tensor change amount and the maximum tensor value of the current computational graph. By comparing the sizes of each tensor in the current computational graph, obtain the maximum tensor size as the maximum tensor value of the current computational graph.

[0104] For example, the quotient of the tensor change divided by the maximum tensor value of the current computational graph can be calculated as the intermediate index value.

[0105] S604. Obtain a criterion value based on the intermediate index value and the simulated temperature of the current computational graph.

[0106] Specifically, a criterion value can be obtained based on the intermediate index value and the simulated temperature of the current computational graph. The simulated temperature of the current computational graph can be obtained by multiplying the preset coefficient by the simulated temperature of the initial computational graph. The simulated temperature of the initial computational graph is obtained in advance, and the preset coefficient is a constant.

[0107] For example, according to the formula calculate the criterion value β, where delta represents the intermediate index value and temp represents the simulated temperature of the current computational graph.

[0108] S605. If the random sampling value is less than the criterion value, update the initial computational graph based on the current computational graph; otherwise, the initial computational graph remains unchanged; where the random sampling value is greater than 0 and less than 1.

[0109] Specifically, compare the criterion value with the random sampling value. If the random sampling value is less than the criterion value, it indicates that the size of the intermediate tensor of the current computational graph is smaller than that of the initial computational graph, which is beneficial to reducing the memory access overhead. Then, the initial computational graph will be updated with the current computational graph, that is, the initial computational graph will be replaced by the current computational graph. If the random sampling value is greater than or equal to the criterion value, it means that the current computational graph does not reduce the intermediate tensor and there is no need to update the initial computational graph, so the initial computational graph will remain unchanged. The random sampling value is a random number between (0, 1).

[0110] S606. Reduce the simulated temperature of the initial computational graph based on a preset coefficient; wherein, the simulated temperature of the first initial computational graph is preset; the preset coefficient is greater than 0 and less than 1;

[0111] Specifically, update the simulated temperature of the initial computational graph with the product result of multiplying the simulated temperature of the initial computational graph by the preset coefficient. The simulated temperature of the first initial computational graph is preset and can be set according to actual needs, which is not limited in the embodiments of the present invention. The preset coefficient is a constant and can be set according to actual needs, which is not limited in the embodiments of the present invention. The preset coefficient is greater than 0 and less than 1.

[0112] For example, the preset coefficient is taken as 0.97, and the simulated temperature of the first initial computational graph is taken as 1000.

[0113] S607. If it is determined that the simulated temperature of the initial computational graph is less than the simulated temperature threshold, it is determined that the initial computational graph meets the preset conditions; otherwise, re-obtain the current computational graph to update the initial computational graph until the equivalent graph of the computational graph is obtained.

[0114] Specifically, compare the simulated temperature of the initial computational graph with the simulated temperature threshold. If the simulated temperature of the initial computational graph is less than the simulated temperature threshold, then it can be determined that the initial computational graph meets the preset conditions, and the initial computational graph is used as the equivalent graph of the computational graph. If the simulated temperature of the initial computational graph is greater than or equal to the simulated temperature threshold, then the current computational graph will be re-obtained, and steps S601 to S607 will be repeated until an initial computational graph that meets the preset conditions is obtained, that is, the equivalent graph of the computational graph is obtained.

[0115] Figure 7 is a schematic flowchart of an optimization method for a tensor program provided by the seventh embodiment of the present invention. As Figure 7 shown, on the basis of the above embodiments, further, the obtaining of multiple computing kernels corresponding to the computational graph by splitting the computing kernel of the equivalent graph of the computational graph includes:

[0116] S701. Obtain the continuous subgraphs of the equivalent graph based on the equivalent graph of the computational graph;

[0117] Specifically, from the equivalent graph of the computation graph, consecutive subgraphs of the equivalent graph can be obtained. A consecutive subgraph includes multiple subgraphs, and the consecutive subgraphs can be obtained by methods in the prior art, which will not be elaborated here.

[0118] S702. Obtain the computational intensity and the number of parallel units of each subgraph in the consecutive subgraphs of the equivalent graph;

[0119] Specifically, for each subgraph in the consecutive subgraphs of the equivalent graph, the computational intensity of the subgraph can be calculated based on the computational amounts of all operators in the subgraph and the tensor sizes of the input tensors and output tensors. The computational intensity is defined as the computational amount of the operator divided by the sum of the sizes of the input tensor and the output tensor. According to the tensor attributes of the input tensors and output tensors in the subgraph, the number of parallel units of the subgraph can be obtained. The number of parallel units of the subgraph is defined as the product result L1 of the dimension lengths corresponding to Batch of the reduction dependency attribute values of all the input tensors and output tensors of the subgraph, and then multiplied by the result L2 of the dimension length corresponding to Reuse of the reduction dependency attribute values of all the input tensors and output tensors of the subgraph divided by a pre-given BlockSize value, that is, L1×L2. Here, the BlockSize value represents the length size covered by one-time computation of each parallel unit in this dimension.

[0120] S703. If it is determined that the computational intensity of the subgraph is greater than or equal to the computational intensity threshold and the number of parallel units of the subgraph is greater than the parallel unit threshold, then use the subgraph as a pending kernel in the set of pending kernels;

[0121] Specifically, compare the computational intensity of the subgraph with the computational intensity threshold, and compare the number of parallel units of the subgraph with the parallel unit threshold. If the computational intensity of the subgraph is greater than or equal to the computational intensity threshold and the number of parallel units of the subgraph is greater than the parallel unit threshold, then use the subgraph as a pending kernel in the set of pending kernels. Among them, the computational intensity threshold is pre-set and set according to actual needs, which is not limited in the embodiments of the present invention. The parallel unit threshold is obtained according to the actual hardware situation, which is not limited in the embodiments of the present invention.

[0122] S704. Based on each pending kernel included in the set of pending kernels and the set screening rule, obtain all combinations of pending kernels;

[0123] Specifically, according to the set screening rule, perform kernel combination screening on each pending kernel included in the set of pending kernels, and various combinations of pending kernels that meet the set screening rule can be obtained. Among them, the set screening rule is preset and set according to actual needs, which is not limited in the embodiments of the present invention.

[0124] For example, the set screening rules include: each pending kernel in the pending kernel combination satisfies the kernel data dependency relationship, and the subgraphs corresponding to the respective pending kernels in the pending kernel combination form a continuous subgraph in the computation graph. That the pending kernel satisfies the kernel data dependency relationship means that the topological order of the pending kernel in the computation graph is consistent with the topological order of the pending kernel in the pending kernel combination.

[0125] Based on the set of pending kernels, all subsets corresponding to the set of pending kernels can be obtained, and each subset includes multiple pending kernels.

[0126] For each subset, if each pending kernel in the subset satisfies the kernel data dependency relationship, and the subgraphs corresponding to the respective pending kernels in the subset are continuous subgraphs in the computation graph, then the subset satisfies the set screening rules, and the subset is taken as a pending kernel combination.

[0127] After traversing all subsets corresponding to the set of pending kernels and obtaining all subsets that satisfy the set screening rules, all pending kernel combinations of the set of pending kernels are obtained.

[0128] S705. Obtain the execution time corresponding to each pending kernel combination based on the execution times corresponding to the respective pending kernels in each pending kernel combination; wherein, the execution time corresponding to a pending kernel is obtained based on the execution code generated for the pending kernel;

[0129] Specifically, for each pending kernel combination, the execution times corresponding to the respective pending kernels in the pending kernel combination can be obtained, and based on the execution times corresponding to the respective pending kernels in the pending kernel combination, the execution time corresponding to the pending kernel combination can be obtained. Wherein, the execution time corresponding to a pending kernel is obtained based on the execution code generated for the pending kernel.

[0130] For example, generate the corresponding Triton code based on the pending kernel, and executing the above Triton code can obtain the execution time of the pending kernel.

[0131] S706. Obtain the pending kernel combination with the least execution time based on the execution times corresponding to all the pending kernel combinations as the multiple computing kernels corresponding to the computation graph.

[0132] Specifically, compare the execution times corresponding to all the pending kernel combinations, obtain the pending kernel combination with the least execution time, take each pending kernel in the pending kernel combination with the least execution time as a computing kernel, and obtain the multiple computing kernels corresponding to the computation graph.

[0133] The technical solution of this application can be applied to the inference optimization of deep learning models. The following takes the inference of a deep learning model in the long sequence text generation scenario as an example to illustrate the specific implementation process of the optimization method of the tensor program provided by the embodiments of the present invention.

[0134] The model used in the scenario of the embodiments of the present invention is H2O. The optimization method of the tensor program provided by the embodiments of the present invention is made into a Python library, providing a Python interface for calling, realizing the optimization of the tensor program and outputting the optimized tensor program code. The core module of H2O is built by Python code.

[0135] Convert the H2O model to the ONNX format. The H2O model can be converted to an ONNX format model through the torch_module_to_onnx conversion function. The torch_module_to_onnx conversion function is pre-written.

[0136] In the embodiments of the present invention, users can directly use a series of function interfaces to optimize the H2O model. The H2O model, as a tensor program, converts the H2O model in the ONNX format into the H2O model in the internal format (module), and then uses the fission and simplify functions to complete the acquisition of the tensor attributes of the computation graph of the H2O model and the acquisition of the equivalent graph. Then, perform computational kernel splitting, use the Connected module and the corresponding optimize method to obtain each pending kernel, use the profile method to actually execute all pending kernels and obtain their execution times, find the optimal kernel combination by calling the codegen method, and generate the optimized Triton code corresponding to the H2O model. The fission function is used to implement the tensor attribute analysis based on the computation graph of the tensor program and the preset tensor attributes in the technical solution of this application, and obtain the tensor attributes of each tensor in the computation graph; the simplify function is used to implement the tensor transformation based on the tensor attributes of each tensor in the computation graph and the transformation rules in the technical solution of this application, and obtain the equivalent graph of the computation graph. The Fission function, simplify function, Connected module, optimize method, and codegen method are all pre-written.

[0137] When the optimized tensor program of this application is executed, it can achieve acceleration in end-to-end performance and core modules compared to other existing technology frameworks. Figure 8 Show the acceleration effects of 7 models, namely H2O, RoCo, Keyformer, SnapKV, Corm, Vanilla Attention, and Gemma2, on the NVIDIA A100 GPU and H100 GPU platforms. FromFigure 8 It can be seen that compared with 6 existing acceleration tools such as PyTorch and TorchInductor, the technical solution of this application can achieve an acceleration effect of about 1.5 times.

[0138] Figure 9 It is a schematic structural diagram of an optimization device for a tensor program provided in the ninth embodiment of the present invention. As Figure 9 shown, the optimization device for a tensor program provided in the embodiments of the present invention includes an analysis module 901, a transformation module 902, and a splitting module 903, where:

[0139] The analysis module 901 is used to perform tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; the transformation module 902 is used to perform tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; where the transformation rules are preset; the splitting module 903 is used to perform computational kernel splitting on the equivalent graph of the computational graph to obtain a plurality of computational kernels corresponding to the computational graph.

[0140] Specifically, the analysis module 901 can obtain the tensor attributes of each tensor in the computational graph by performing tensor attribute analysis on the computational graph of the tensor program according to the preset tensor attributes. The preset tensor attributes are tensor attributes that are preset and play an important role in the optimization of the tensor program, and are set according to actual needs, which are not limited in the embodiments of the present invention.

[0141] After obtaining the tensor attributes of all tensors in the computational graph of the tensor program, the transformation module 902 can perform tensor transformation according to the tensor attributes of each tensor in the computational graph and the transformation rules, and use the transformed tensors to replace the corresponding tensors in the computational graph to obtain an equivalent graph of the computational graph. The transformation of tensors mainly considers the impact of the size of intermediate tensors in the computational graph on memory access. Through the transformation rules, the size of intermediate tensors can be reduced, thereby reducing the memory access overhead and improving the overall performance. The transformation rules are preset and are set according to actual needs, which are not limited in the embodiments of the present invention.

[0142] After obtaining the equivalent graph of the computational graph, the splitting module 903 performs computational kernel splitting on the equivalent graph of the computational graph, and can split the equivalent graph of the computational graph into a plurality of computational kernels. Each computational kernel will be handed over to the GPU to execute the corresponding calculation. The above-mentioned plurality of computational kernels can be a combination including non-convex kernels, and non-convex kernels have better parallelism and square efficiency, which can improve the execution efficiency of the tensor program.

[0143] The optimization method and device for tensor programs provided by the embodiments of the present invention can perform tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; perform tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; and perform computational kernel splitting on the equivalent graph of the computational graph to obtain multiple computational kernels corresponding to the computational graph, which can improve the execution efficiency of the tensor program.

[0144] Based on the above embodiments, further, the analysis module 901 is specifically configured to:

[0145] If the computational graph of the tensor program includes a single operator, obtain the tensor attributes of the input tensor and output tensor of the single operator based on the computational semantics of the single operator and preset tensor attributes.

[0146] Based on the above embodiments, further, the analysis module 901 is specifically configured to:

[0147] If the computational graph of the tensor program includes multiple operators, perform iterative calculation on each operator included in the computational graph until there is no situation where a high-priority value of the same tensor attribute replaces a low-priority value in the input tensors and output tensors of all operators included in the computational graph.

[0148] Figure 10 It is a schematic structural diagram of the optimization device for tensor programs provided by the tenth embodiment of the present invention. As Figure 10 shown, based on the above embodiments, further, the transformation module 902 includes a traversal unit 9021 and an update unit 9022, where:

[0149] The traversal unit 9021 is configured to traverse each operator in the computational graph, and obtain the operators that meet the transformation rules and the transformed expressions corresponding to each operator based on the input tensors and output tensors of each operator and the transformation rules; the update unit 9022 is configured to update the computational graph based on the operators that meet the transformation rules and their respective transformed expressions until a computational graph that meets the preset conditions is obtained as the equivalent graph of the computational graph.

[0150] Figure 11 It is a schematic structural diagram of the optimization device for tensor programs provided by the eleventh embodiment of the present invention. As Figure 11 shown, based on the above embodiments, further, the traversal unit 9021 includes a first obtaining subunit 90211, a first judging subunit 90212, a second judging subunit 90213, and an acting subunit 90214, where:

[0151] The first obtaining subunit 90211 is configured to obtain the transformed expression of the operator under each matching transformation condition based on the input tensor of the operator and the transformation rule; wherein, the transformation rule includes at least one transformation condition, and each transformation condition has a corresponding legality judgment condition; the first judgment subunit 90212 is configured to, if the sum of the sizes of the transformed input tensor and output tensor of the operator under the matching transformation condition is less than the sum of the sizes of the transformed input tensor and output tensor of the operator under the matching transformation condition and the legality judgment condition corresponding to the transformation condition matched by the operator is empty, then regard the operator as an operator that satisfies the transformation rule; the second judgment subunit 90213 is configured to, if the sum of the sizes of the transformed input tensor and output tensor of the operator under the transformation condition is less than the sum of the sizes of the input tensor and output tensor of the operator before transformation and the legality judgment condition corresponding to the transformation condition corresponding to the operator is not empty, then judge whether the transformation of the operator under the transformation condition is legal based on the legality judgment condition corresponding to the transformation condition corresponding to the operator and the tensor attribute of the input tensor of the operator; the making subunit 90214 is configured to, if the transformation of the operator under the transformation condition is legal, then regard the operator as an operator that satisfies the transformation rule.

[0152] Figure 12 It is a schematic structural diagram of an optimization device for a tensor program provided in the twelfth embodiment of the present invention. As Figure 12 shown, on the basis of the above embodiments, further, the update unit 9022 includes a second obtaining subunit 90221, a calculating subunit 90222, a second obtaining subunit 90223, a third obtaining subunit 90224, a third judgment subunit 90225, a reducing subunit 90226, and a fourth judgment subunit 90227, where:

[0153] The second acquisition subunit 90221 is configured to acquire the transformed expression corresponding to at least one operator that satisfies the transformation rule, update the computation graph, and obtain the current computation graph; the computation subunit 90222 is configured to calculate the tensor size of the initial computation graph minus the tensor size of the current computation graph to obtain a tensor change amount; wherein, the first initial computation graph uses the computation graph; the second acquisition subunit 90223 is configured to obtain an intermediate index value based on the tensor change amount and the maximum tensor value of the current computation graph; the third acquisition subunit 90224 is configured to obtain a criterion value according to the intermediate index value and the simulated temperature of the current computation graph; the third determination subunit 90225 is configured to, if the random sampling value is less than the criterion value, update the initial computation graph based on the current computation graph; otherwise, the initial computation graph remains unchanged; wherein, the random sampling value is greater than 0 and less than 1; the reduction subunit 90226 is configured to reduce the simulated temperature of the initial computation graph based on a preset coefficient; wherein, the simulated temperature of the first initial computation graph is preset; the preset coefficient is greater than 0 and less than 1; the fourth determination subunit 90227 is configured to, if it is determined that the simulated temperature of the initial computation graph is less than the simulated temperature threshold, determine that the initial computation graph meets the preset condition; otherwise, re-acquire the current computation graph to update the initial computation graph until an equivalent graph of the computation graph is obtained.

[0154] Figure 13 It is a schematic structural diagram of an optimization device for a tensor program provided in the thirteenth embodiment of the present invention, as Figure 13 shown. On the basis of the above embodiments, further, the splitting module 903 includes a first acquisition unit 9031, an acquisition unit 9032, a determination unit 9033, a second acquisition unit 9034, a third acquisition unit 9035, and a fourth acquisition unit 9036, wherein:

[0155] The first acquisition unit 9031 is configured to obtain a continuous subgraph of the equivalent graph based on the equivalent graph of the computation graph; the acquisition unit 9032 is configured to acquire the computation intensity and the number of parallel units of each subgraph in the continuous subgraph of the equivalent graph; the determination unit 9033 is configured to, if it is determined that the computation intensity of the subgraph is greater than or equal to the computation intensity threshold and the number of parallel units of the subgraph is greater than the parallel unit threshold, use the subgraph as a pending kernel in the pending kernel set; the second acquisition unit 9034 is configured to obtain all pending kernel combinations based on each pending kernel included in the pending kernel set and the set screening rule; the third acquisition unit 9035 is configured to obtain the execution time corresponding to each pending kernel combination based on the execution time corresponding to each pending kernel in each pending kernel combination; wherein, the execution time corresponding to the pending kernel is obtained based on the execution code generated by the pending kernel; the fourth acquisition unit 9036 is configured to obtain the pending kernel combination with the least execution time as the multiple computation kernels corresponding to the computation graph.

[0156] The embodiments of the device provided by the embodiments of the present invention can specifically be used to execute the processing procedures of the above method embodiments, and their functions will not be elaborated here. Reference can be made to the detailed descriptions of the above method embodiments.

[0157] Figure 14 is a schematic physical structure diagram of a computer device provided by the eleventh embodiment of the present invention. As Figure 14 shown, the computer device 600 may include: a processor 100 and a memory 140. The memory 140 is coupled to the processor 100. The processor 100 can call the logical instructions in the memory 140 to execute the methods provided by the above method embodiments, for example, including: performing tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; performing tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; where the transformation rules are preset; performing computational kernel splitting on the equivalent graph of the computational graph to obtain a plurality of computational kernels corresponding to the computational graph.

[0158] This embodiment discloses a computer program product. The computer program product includes computer programs / instructions stored on a computer-readable storage medium. When the computer programs / instructions are executed by a computer, the computer can execute the methods provided by the above method embodiments, for example, including: performing tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; performing tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; where the transformation rules are preset; performing computational kernel splitting on the equivalent graph of the computational graph to obtain a plurality of computational kernels corresponding to the computational graph.

[0159] This embodiment provides a computer-readable storage medium. The computer-readable storage medium stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the computer executes the methods provided by the above method embodiments, for example, including: performing tensor attribute analysis based on the computational graph of the tensor program and preset tensor attributes to obtain the tensor attributes of each tensor in the computational graph; performing tensor transformation based on the tensor attributes of each tensor in the computational graph and transformation rules to obtain an equivalent graph of the computational graph; where the transformation rules are preset; performing computational kernel splitting on the equivalent graph of the computational graph to obtain a plurality of computational kernels corresponding to the computational graph.

[0160] As Figure 14As shown, the computer device 600 may further include: a communication module 110, an input unit 120, an audio processor 130, a display 160, and a power supply 170. It should be noted that the computer device 600 does not necessarily have to include Figure 14 all the components shown therein; in addition, the computer device 600 may further include Figure 13 components not shown in the figure. Reference may be made to the prior art. It should be noted that this figure is exemplary; other types of structures may also be used to supplement or replace this structure to achieve telecommunications functions or other functions.

[0161] As Figure 14 shown, the processor 100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The processor 100 receives inputs and controls the operations of the various components of the computer device 600.

[0162] Among them, the memory 140 may be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the processor 100 can execute the program stored in the memory 140 to implement information storage or processing, etc.

[0163] The input unit 120 provides inputs to the processor 100. The input unit 120 is, for example, a key or a touch input device. The power supply 170 is used to supply power to the computer device 600. The display 160 is used to display display objects such as images and texts. The display 160 may be, for example, an LCD display, but is not limited thereto.

[0164] The memory 140 may be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be such a memory that stores information even when powered off, can be selectively erased and has more data. Examples of the memory 140 are sometimes referred to as EPROMs, etc. The memory 140 may also be some other type of device. The memory 140 includes a buffer 141 (sometimes referred to as a buffer memory). The memory 140 may include an application / function storage unit 142, and the application / function storage unit 142 is used to store application programs and function programs or the processes for operating the computer device 600 through the processor 100.

[0165] The memory 140 may further include a data storage unit 143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the computer device. The driver storage unit 144 of the memory 140 may include various drivers of the computer device for communication functions and / or for performing other functions of the computer device (such as a messaging application, an address book application, etc.).

[0166] The communication module 110 includes a transmitter / receiver that transmits and receives signals via the antenna 111. The communication module 110 is coupled to the processor 100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.

[0167] Based on different communication technologies, multiple communication modules 110 may be provided in the same computer device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 110 is also coupled to the speaker 131 and the microphone 132 via the audio processor 130 to provide an audio output via the speaker 131 and receive an audio input from the microphone 132, thereby implementing normal telecommunication functions. The audio processor 130 may include any suitable buffers, decoders, amplifiers, etc. Additionally, the audio processor 130 is also coupled to the processor 100, so that recording can be performed on the local machine through the microphone 132, and the sound stored on the local machine can be played through the speaker 131.

[0168] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0169] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the acts Figure 1 or acts and / or boxes Figure 1 specified in one or more of the boxes.

[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the acts Figure 1 or acts and / or boxes Figure 1 specified in one or more of the boxes.

[0172] In the description of the present specification, the descriptions with reference to the terms "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In the present specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0173] The above-described specific embodiments have further elaborated on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing a tensor program, characterized in that: include: Performing tensor attribute analysis based on a computation graph of a tensor program and preset tensor attributes to obtain tensor attributes of each tensor in the computation graph; Performing tensor transformation based on tensor attributes and transformation rules of each tensor in the computation graph to obtain an equivalent graph of the computation graph; wherein the transformation rules are preset; The equivalent graph of the computation graph is split into computation kernels to obtain multiple computation kernels corresponding to the computation graph.

2. The method according to claim 1, characterized in that The tensor attribute analysis is performed based on the calculation graph of the tensor program and the preset tensor attributes to obtain the tensor attributes of each tensor in the calculation graph, including: If the computation graph of the tensor program includes a single operator, tensor attributes of an input tensor and an output tensor of the single operator are obtained based on the computation semantics and preset tensor attributes of the single operator.

3. The method according to claim 1, characterized in that The tensor attribute analysis is performed based on the calculation graph of the tensor program and the preset tensor attributes to obtain the tensor attributes of each tensor in the calculation graph, including: If the computation graph of the tensor program includes multiple operators, iterative calculation is performed on each operator included in the computation graph until the input tensors and output tensors of all operators included in the computation graph do not have a situation where a high-priority value replaces a low-priority value of the same tensor attribute.

4. The method according to claim 1, characterized in that: The performing tensor transformation based on the tensor attributes and transformation rules of each tensor in the computation graph to obtain an equivalent graph of the computation graph includes: Traversing each operator in the computation graph, and based on the input tensor and output tensor of each operator and the transformation rule, obtaining operators satisfying the transformation rule and transformed expressions corresponding to each operator; The computation graph is updated based on each operator satisfying the transformation rule and the corresponding transformed expressions, until a computation graph satisfying a preset condition is obtained as an equivalent graph of the computation graph.

5. The method according to claim 4, characterized in that The step of obtaining the operators satisfying the transformation rules and the transformed expressions corresponding to the operators based on the input tensors and output tensors of each operator and the transformation rules includes: Based on the input tensor of the operator and the transformation rule, obtaining the transformed expression of the operator under each matched transformation condition; wherein the transformation rule includes at least one transformation condition, and each transformation condition has a corresponding legality judgment condition; If the sum of the sizes of the transformed input tensor and the output tensor of the operator under the matched transformation condition is smaller than the sum of the sizes of the transformed input tensor and the output tensor of the operator under the matched transformation condition and the legality judgment condition corresponding to the transformation condition matched by the operator is empty, then the operator is regarded as an operator that satisfies the transformation rule; If the sum of the sizes of the input tensor and the output tensor of the operator after transformation under the transformation condition is less than the sum of the sizes of the input tensor and the output tensor of the operator before transformation and the legality judgment condition corresponding to the transformation condition corresponding to the operator is not empty, then based on the legality judgment condition corresponding to the transformation condition corresponding to the operator and the tensor attribute of the input tensor of the operator, it is judged whether the transformation of the operator under the transformation condition is legal; If the transformation of the operator under the transformation condition is legal, the operator is regarded as an operator satisfying the transformation rule.

6. The method according to claim 4, characterized in that The updating of the computation graph based on the operators satisfying the transformation rule and the transformed expressions corresponding to the operators until a computation graph satisfying a preset condition is obtained as an equivalent graph of the computation graph includes: Obtaining at least one transformed expression corresponding to an operator satisfying the transformation rule to update the computation graph, and obtaining a current computation graph; Calculate the tensor size of the initial calculation graph minus the tensor size of the current calculation graph to obtain the tensor change; wherein the first initial calculation graph adopts the calculation graph; Based on the tensor change amount and the maximum tensor value of the current calculation graph, an intermediate index value is obtained; Obtaining a criterion value according to the intermediate index value and the simulated temperature of the current calculation diagram; If the random sampling value is less than the criterion value, the initial calculation graph is updated based on the current calculation graph; otherwise, the initial calculation graph remains unchanged; wherein the random sampling value is greater than 0 and less than 1; Lowering the simulation temperature of the initial calculation graph based on a preset coefficient; wherein the simulation temperature of the first initial calculation graph is preset; and the preset coefficient is greater than 0 and less than 1; If it is determined that the simulation temperature of the initial calculation graph is less than the simulation temperature threshold, it is determined that the initial calculation graph meets the preset conditions; otherwise, the current calculation graph is re-acquired to update the initial calculation graph until an equivalent graph of the calculation graph is obtained.

7. The method according to any one of claims 1 to 6, characterized in that: The step of performing computational kernel splitting on the equivalent graph of the computational graph to obtain a plurality of computational kernels corresponding to the computational graph comprises: Based on the equivalent graph of the computation graph, obtaining a continuous subgraph of the equivalent graph; Obtaining the computational intensity and the number of parallel units of each subgraph in the continuous subgraph of the equivalent graph; If it is determined that the computational intensity of the subgraph is greater than or equal to the computational intensity threshold and the number of parallel units of the subgraph is greater than the parallel unit threshold, the subgraph is used as a pending kernel in the pending kernel set; Based on the individual pending kernels included in the pending kernel set and the set screening rule, all pending kernel combinations are obtained; Based on the execution time corresponding to each pending kernel of each pending kernel combination, the execution time corresponding to each pending kernel combination is obtained; wherein the execution time corresponding to the pending kernel is obtained based on the execution code generated by the pending kernel; Based on the execution times corresponding to all pending kernel combinations, a pending kernel combination with the shortest execution time is obtained as the multiple computing kernels corresponding to the computing graph.

8. A tensor program optimization device, characterized in that: include: An analysis module, used to perform tensor attribute analysis based on a computation graph of a tensor program and preset tensor attributes, and obtain tensor attributes of each tensor in the computation graph; A transformation module, used for performing tensor transformation based on tensor attributes and transformation rules of each tensor in the computational graph to obtain an equivalent graph of the computational graph; wherein the transformation rules are preset; The splitting module is used to split the equivalent graph of the computational graph into computational kernels to obtain multiple computational kernels corresponding to the computational graph.

9. The device according to claim 8, characterized in that The analysis module is specifically used for: If the computation graph of the tensor program includes a single operator, tensor attributes of an input tensor and an output tensor of the single operator are obtained based on the computation semantics and preset tensor attributes of the single operator.

10. The device according to claim 8, characterized in that The analysis module is specifically used for: If the computation graph of the tensor program includes multiple operators, iterative calculation is performed on each operator included in the computation graph until the input tensors and output tensors of all operators included in the computation graph do not have a situation where a high-priority value replaces a low-priority value of the same tensor attribute.

11. The device according to claim 8, characterized in that The transformation module comprises: A traversal unit, used to traverse each operator in the computational graph, and based on the input tensor and output tensor of each operator and the transformation rule, obtain the operator satisfying the transformation rule and the transformed expression corresponding to each operator; An updating unit is used to update the computational graph based on each operator satisfying the transformation rule and the corresponding transformed expressions, until a computational graph satisfying a preset condition is obtained as an equivalent graph of the computational graph.

12. The device according to claim 11, characterized in that The traversal unit comprises: A first obtaining subunit, configured to obtain a transformed expression of the operator under each matched transformation condition based on the input tensor of the operator and the transformation rule; wherein the transformation rule includes at least one transformation condition, and each transformation condition has a corresponding legality judgment condition; A first judgment subunit, configured to treat the operator as an operator that satisfies the transformation rule if the sum of the sizes of the transformed input tensor and the output tensor of the operator under the matched transformation condition is less than the sum of the sizes of the transformed input tensor and the output tensor of the operator under the matched transformation condition and the legality judgment condition corresponding to the transformation condition matched by the operator is empty; A second judgment subunit is used to judge whether the transformation of the operator under the transformation condition is legal based on the legality judgment condition corresponding to the transformation condition corresponding to the operator and the tensor attribute of the input tensor of the operator, if the sum of the sizes of the input tensor and the output tensor of the operator after the transformation under the transformation condition is less than the sum of the sizes of the input tensor and the output tensor of the operator before the transformation and the legality judgment condition corresponding to the transformation condition corresponding to the operator is not empty; As a subunit, if the transformation of the operator under the transformation condition is legal, then the operator is regarded as an operator that satisfies the transformation rule.

13. The device according to claim 11, characterized in that The updating unit comprises: A second acquisition subunit is used to acquire at least one transformed expression corresponding to an operator that satisfies the transformation rule to update the calculation graph and obtain a current calculation graph; A computing subunit, used to calculate the tensor size of the initial computing graph minus the tensor size of the current computing graph, to obtain a tensor change; wherein the first initial computing graph adopts the computing graph; A second obtaining subunit, used to obtain an intermediate index value based on the tensor change amount and the maximum tensor value of the current calculation graph; A third obtaining subunit is used to obtain a criterion value according to the intermediate index value and the simulated temperature of the current calculation diagram; A third judgment subunit is used to update the initial calculation graph based on the current calculation graph if the random sampling value is less than the criterion value; otherwise, the initial calculation graph remains unchanged; wherein the random sampling value is greater than 0 and less than 1; A reducing subunit, used to reduce the simulation temperature of the initial calculation graph based on a preset coefficient; wherein the simulation temperature of the first initial calculation graph is preset; and the preset coefficient is greater than 0 and less than 1; The fourth judgment subunit is used to determine that the initial calculation graph meets the preset conditions if it is determined that the simulation temperature of the initial calculation graph is less than the simulation temperature threshold; otherwise, re-acquire the current calculation graph to update the initial calculation graph until an equivalent graph of the calculation graph is obtained.

14. The device according to any one of claims 8 to 13, characterized in that The splitting module comprises: A first obtaining unit, configured to obtain a continuous subgraph of the equivalent graph based on the equivalent graph of the computation graph; An acquisition unit, used for acquiring the computational intensity and the number of parallel units of each subgraph in the continuous subgraph of the equivalent graph; a judging unit, configured to take the subgraph as a pending kernel in a pending kernel set if it is judged that the computing intensity of the subgraph is greater than or equal to a computing intensity threshold and the number of parallel units of the subgraph is greater than a parallel unit threshold; A second obtaining unit, configured to obtain all pending kernel combinations based on the pending kernels included in the pending kernel set and a set screening rule; A third obtaining unit is used to obtain the execution time corresponding to each pending kernel combination based on the execution time corresponding to each pending kernel of each pending kernel combination; wherein the execution time corresponding to the pending kernel is obtained based on the execution code generated by the pending kernel; The fourth obtaining unit is used to obtain, based on the execution times corresponding to all the pending kernel combinations, a pending kernel combination with the shortest execution time as the multiple computing kernels corresponding to the computing graph.

15. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

17. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.