Computation graph splitting method and related equipment
Patent Information
- Application Number
- CN202380072302.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-05-27
AI Technical Summary
In scenarios such as target classification and speech recognition, the amount of calculations in the calculation graph is unbalanced, making it difficult to reasonably split existing technologies and affecting processing performance.
Obtain the dependencies between operators through polyhedral modeling, split the calculation graph based on these relationships, and optimize the subgraphs to improve the processing performance of data processing equipment.
It achieves a more reasonable split of the calculation graph, maximizes the data locality and parallelism of the subgraphs, improves the accuracy of processing performance and fusion analysis, and can complete tuning in a short time to reach industrial usable levels.
Smart Images

Figure CN120051772A_ABST
Abstract
Description
A computational graph splitting method and related equipment Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a computational graph splitting method and related equipment. Background Art
[0002] Artificial intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and basic AI theory.
[0003] Currently, computational graphs are often used to express network structures in scenarios such as object classification and speech recognition. A computational graph is a directed graph with operators as nodes, representing computational functions.
[0004] However, in practical applications, the amount of computation carried by each operator often varies, and the computational complexity also varies. Therefore, how to properly split the computation graph is a technical problem that needs to be solved urgently.
[0005] Summary of the Invention
[0006] The present invention provides a computational graph splitting method and related devices. During the computational graph splitting process, the dependency relationships between multiple statements obtained by polyhedron modeling can be considered to make the computational graph splitting more reasonable.
[0007] In a first aspect, an embodiment of the present application provides a method for splitting a computation graph, which can be performed by a data processing device or a component of the data processing device (e.g., a processor, a chip, or a chip system). The method includes: obtaining a computation graph, where the computation graph is used to represent the computational logic of multiple operator nodes; performing polyhedral modeling on the computation graph to obtain multiple statements; and splitting the computation graph based on the types of dependency relationships between the multiple statements to obtain multiple first subgraphs, where the dependency relationships are used to represent the sequential relationships of reads and / or writes between the multiple statements.
[0008] In an embodiment of the present application, multiple statements are obtained by polyhedron modeling of the computation graph, and then the computation graph is split according to the types of dependency relationships between the multiple statements. Because the dependency relationships obtained based on polyhedron modeling are based on the granularity of loops in operators, compared to the existing solutions that only use operators as the granularity, this method can make the computation graph splitting more reasonable by considering the dependency relationships between the multiple statements obtained by polyhedron modeling during the computation graph splitting process.
[0009] Optionally, in a possible implementation of the first aspect, the above method is applied to a data processing device, and the above method also includes: optimizing multiple first subgraphs to obtain multiple second subgraphs, the number of the multiple second subgraphs is greater than or equal to the number of the multiple first subgraphs, and the multiple second subgraphs are used to improve the processing performance of the data processing device.
[0010] In this possible implementation, compared to existing methods that suffer from search space explosion and are unable to achieve subgraph tuning, this method leverages a polyhedron modeling subgraph partitioning strategy to maximize subgraph data locality and parallelism, thereby improving the accuracy of fusion analysis. It then further optimizes multiple first subgraphs to generate multiple second subgraphs. This allows for rapid tuning, achieving industrial-grade performance.
[0011] Optionally, in a possible implementation of the first aspect, the above-mentioned types include at least one of the following: a first type, a second type, a third type, a fourth type, a fifth type and a sixth type; the first type is used to indicate that the access order between the statements is the same, and the number of instances of write operations and read operations between the statements is the same; the second type is used to indicate that the access order between the statements is different, and the number of instances of write operations and read operations between the statements is the same; the third type is used to indicate that the number of write operations of each statement is greater than the number of read operations; the fourth type is used to indicate that the number of write operations of each statement is less than the number of read operations; the fifth type is used to indicate that the mapping relationship between the statements is one-to-one; and the sixth type is used to indicate that the mapping relationship between the statements is not one-to-one.
[0012] In this possible implementation, dependencies are classified by information related to read and write operations (such as memory, order, etc.), so that the computational graph can be split according to the information related to read and write operations to improve the utilization of computing units in data processing equipment.
[0013] Optionally, in a possible implementation of the first aspect, the above steps of: splitting the computation graph based on the type to obtain multiple first subgraphs, include: determining the fusion level between statements corresponding to the type based on the mapping relationship, the mapping relationship is used to describe the correspondence between the type and the fusion level, and the fusion level is used to describe the fusibility of each statement; splitting the computation graph based on the fusion level to obtain multiple first subgraphs.
[0014] In this possible implementation, the fusibility of the dependency analysis operators obtained from the polyhedron model is analyzed. The subgraph segmentation strategy based on polyhedron modeling can maximize the data locality and parallelism of the subgraph, thereby improving the accuracy of the fusibility analysis.
[0015] Optionally, in a possible implementation of the first aspect, the above-mentioned fusion level includes at least one of the following: the first level, the second level, the third level and the fourth level; the first level is used to indicate that each sentence can be fused front to back, the second level is used to indicate that each sentence cannot be fused back to back, the third level is used to indicate that each sentence cannot be fused front to back, and the fourth level is used to indicate that each sentence can be fused front to back.
[0016] In this possible implementation, by combining the ability of polyhedron model scheduling and using the possibility of front-end fusion to divide the fusion level, a more direct fusion between statements can be obtained.
[0017] Optionally, in a possible implementation of the first aspect, the above-mentioned mapping relationships include: the correspondence between the first level and the first type, the correspondence between the first level and the fifth type, the correspondence between the second level and the fourth type, the correspondence between the third level and the first combination type, the correspondence between the fourth level and the second type, the correspondence between the fourth level and the third type, and the correspondence between the fourth level and the sixth type; the first combination type includes the combination of the second type and the fourth type.
[0018] In this possible implementation, several specific examples of dependency types and fusion levels are provided to improve the feasibility of the solution.
[0019] A second aspect of an embodiment of the present application provides a data processing device, which includes: an acquisition unit for acquiring a computational graph, the computational graph being used to represent the computational logic of multiple operator nodes; a modeling unit for performing polyhedron modeling on the computational graph to obtain multiple statements; and a splitting unit for splitting the computational graph based on the type of dependency relationship between the multiple statements to obtain multiple first subgraphs, the dependency relationship being used to represent the sequential relationship of reading and / or writing between the multiple statements.
[0020] Optionally, in a possible implementation of the second aspect, the above-mentioned data processing device also includes: an optimization unit, used to optimize multiple first subgraphs to obtain multiple second subgraphs, the number of the multiple second subgraphs is greater than or equal to the number of the multiple first subgraphs, and the multiple second subgraphs are used to improve the processing performance of the data processing device.
[0021] Optionally, in a possible implementation of the second aspect, the above-mentioned types include at least one of the following: a first type, a second type, a third type, a fourth type, a fifth type and a sixth type; the first type is used to indicate that the access order between the statements is the same, and the number of instances of write operations and read operations between the statements is the same; the second type is used to indicate that the access order between the statements is different, and the number of instances of write operations and read operations between the statements is the same; the third type is used to indicate that the number of write operations of each statement is greater than the number of read operations; the fourth type is used to indicate that the number of write operations of each statement is less than the number of read operations; the fifth type is used to indicate that the mapping relationship between the statements is one-to-one; and the sixth type is used to indicate that the mapping relationship between the statements is not one-to-one.
[0022] Optionally, in a possible implementation of the second aspect, the above-mentioned splitting unit is specifically used to determine the fusion level between statements with corresponding types based on a mapping relationship, the mapping relationship is used to describe the correspondence between the type and the fusion level, and the fusion level is used to describe the fusibility of each statement; the splitting unit is specifically used to split the computational graph based on the fusion level to obtain multiple first subgraphs.
[0023] Optionally, in a possible implementation of the second aspect, the above-mentioned fusion level includes at least one of the following: the first level, the second level, the third level and the fourth level; the first level is used to indicate that each sentence can be fused front to back, the second level is used to indicate that each sentence cannot be fused back to back, the third level is used to indicate that each sentence cannot be fused front to back, and the fourth level is used to indicate that each sentence can be fused front to back.
[0024] Optionally, in a possible implementation of the second aspect, the above-mentioned mapping relationships include: the correspondence between the first level and the first type, the correspondence between the first level and the fifth type, the correspondence between the second level and the fourth type, the correspondence between the third level and the first combination type, the correspondence between the fourth level and the second type, the correspondence between the fourth level and the third type, and the correspondence between the fourth level and the sixth type; the first combination type includes the combination of the second type and the fourth type.
[0025] A third aspect of an embodiment of the present application provides a data processing device, including: a processor, the processor is coupled to a memory, the memory is used to store programs or instructions, and when the program or instructions are executed by the processor, the data processing device implements the method in the above-mentioned first aspect or any possible implementation of the first aspect.
[0026] A fourth aspect of an embodiment of the present application provides a computer-readable medium having a computer program or instruction stored thereon. When the computer program or instruction runs on a computer, the computer executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0027] A fifth aspect of the embodiments of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0028] Among them, the technical effects brought about by the second, third, fourth, and fifth aspects or any possible implementation methods thereof can refer to the technical effects brought about by the first aspect or different possible implementation methods of the first aspect, and will not be repeated here.
[0029] It can be seen from the above technical solution that the embodiment of the present application has the following advantages: by polyhedron modeling of the computation graph to obtain multiple statements, and then splitting the computation graph according to the type of dependency relationship between the multiple statements, in the process of splitting the computation graph, the dependency relationship between the multiple statements obtained by polyhedron modeling is considered, so that the splitting of the computation graph can be made more reasonable. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1 is a spatial schematic diagram of a statement provided in an embodiment of the present application;
[0031] FIG2 is a schematic diagram of the structure of the system architecture provided in an embodiment of the present application;
[0032] FIG3 is a flow chart of a computation graph splitting method provided in an embodiment of the present application;
[0033] FIG4 is a diagram illustrating a structure of a computational graph provided in an embodiment of the present application;
[0034] FIG5 is an example diagram of a Json script provided in an embodiment of the present application;
[0035] FIG6 is an example diagram of the first type of dependency relationship provided in an embodiment of the present application;
[0036] FIG7 is an example diagram of the second type of dependency relationship provided in an embodiment of the present application;
[0037] FIG8 is another example diagram of the second type of dependency relationship provided in an embodiment of the present application;
[0038] FIG9 is an example diagram of the third type of dependency relationship provided in an embodiment of the present application;
[0039] FIG10 is an example diagram of the fourth type of dependency relationship provided in an embodiment of the present application;
[0040] FIG11 is a schematic diagram of a flow chart of a subgraph segmentation method provided in an embodiment of the present application;
[0041] FIG12 is a schematic diagram of a flow chart of a subgraph optimization method provided in an embodiment of the present application;
[0042] FIG13 is an example diagram of a subgraph segmentation scheme selected by Task2 according to an embodiment of the present application;
[0043] FIG14 is an example diagram of a subgraph segmentation scheme selected by Task3 according to an embodiment of the present application;
[0044] FIG15 is a schematic diagram showing a performance comparison between the method provided in an embodiment of the present application and the prior art in a training scenario;
[0045] FIG16 is a schematic diagram showing a performance comparison between the method provided in an embodiment of the present application and the prior art in an inference scenario;
[0046] FIG17 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0047] FIG18 is another structural diagram of the data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The present invention provides a computational graph splitting method and related devices. During the computational graph splitting process, the dependency relationships between multiple statements obtained by polyhedron modeling can be considered to make the computational graph splitting more reasonable.
[0049] To facilitate understanding, the following first introduces the relevant terms and concepts mainly involved in the embodiments of this application.
[0050] 1. Artificial intelligence (AI) framework
[0051] AI frameworks typically use a multi-level IR design, primarily to meet both ease-of-use and high-performance requirements. For example, to make it easier for developers to use, the front-end (layer) of an AI framework typically uses a high-level IR. This IR abstracts and encapsulates Tensor computations as much as possible, allowing developers to focus on logically defined models and operators. On the other hand, when optimizing operator performance at the back-end (operator layer) of the AI framework, a low-level IR is typically used. This IR allows the operator layer compiler to combine different hardware characteristics to perform more fine-grained optimizations. The following is a brief introduction to mainstream layer and operator layer compilers:
[0052] Mainstream graph compilers include Mindspore's MindCompiler, Tensorflow's XLA, and TVM's Relay, all of which focus on non-loop optimizations. In addition to common optimizations in traditional compilers, such as constant folding, algebraic simplification, and common subexpressions, they also perform layout conversion and operator fusion (related to the embodiments of this application). By analyzing and optimizing the existing network computational graph logic, they split, reorganize, and fuse the original computational logic to reduce the overhead between operator executions and improve device computing resource utilization, thereby optimizing the overall network execution time.
[0053] Mainstream operator compilers include MindSpore AKG (related to the embodiments of this application), CANN TBE, and TVM. Low-level IR optimizations primarily include scheduling-related optimizations such as loop transformation and loop splitting, as well as back-end pass optimizations such as hardware intrinsic mapping and memory allocation.
[0054] 2. Halide Intermediate Language (IR)
[0055] HalideIR is a common language used to develop high-performance image processing and array computing. It can be used as an intermediate expression to decouple algorithms and optimizations.
[0056] 3. Poly statement (referred to as statement)
[0057] The Poly statement is the basic unit of scheduling in the polyhedron model, represented by S_index, for example: S_0, S_1. The above example is described in detail below. An example of a loop nesting is as follows:
[0058] for(int i=1;i <N;i++)
[0059] for(int j=1; j <N;j++)
[0060] A[i,j]=f(A[i-1][j],A[i][j-1]);
[0061] Where N is a constant. Statements within the loop nest update the data at position A[i][j] by referencing the data stored in A[i-1][j] and A[i][j-1]. If each iteration of the statement within the loop is abstracted as a point in space, a two-dimensional space based on (i, j) can be constructed (as shown in Figure 1). Each black point in Figure 1 represents an iteration of the statement that writes to A[i][j]. This allows us to construct a rectangle consisting of all the black points. This rectangle can be considered a polyhedron in two-dimensional space, and this space is called the iteration space of the computation. Furthermore, this two-dimensional polyhedron can be represented using a set from algebra: {[i, j]: 1 <= i <= N-1 and 1 <= j <= N–1}, where [i, j] is a two-tuple and the inequality following the “:” represents the interval of this set. Let this tuple be called S, representing a statement, then the Polyhedron of this statement can be expressed as {S[i, j]: 1<=i<=N-1 and 1<=j<=N–1}.
[0062] 4. Global tensor
[0063] The global tensor is the tensor used as the external input of the program, rather than the temporary tensor (intermediate tensor) generated in the middle of the program.
[0064] 5. Aggregate subgraph
[0065] An aggregated subgraph represents a large computational graph in which different operators are expanded into basic operators and fused together in the graph-computation fusion feature.
[0066] 6. Computational Graph
[0067] AI frameworks, including Mindspore, all use computation graphs to represent network structures. A computation graph is a directed graph with operators as nodes, representing a computation function. Within the AI framework, this computation function sequentially executes the operator nodes in the directed graph on the input tensor, producing the final output tensor. Operators are essentially computation functions. For example, Add and Sigmoid are both called operators. However, a complex operator like Sigmoid is mathematically implemented using basic operators such as Neg, Exp, Add, and Reciprocal. Therefore, the above analysis shows that computation graphs and operators are essentially the same in terms of computational nature. Operators are packaged computation graphs, while computation graphs are unpacked operators. Therefore, in theory, one can first define a small set of "basic operators" and then use one or more basic operators to equivalently represent any existing operator, thereby further expressing any existing computation graph. This computation graph composed of multiple basic operators is called a "fused operator." Alternatively, the fusion subgraph / fusion operator refers to all subgraphs that contain different basic operators. From the layer level, it is a subgraph, and from the operator level, it is a fusion operator.
[0068] 7. MicroGraph
[0069] MicroGraph is a subgraph obtained by splitting an aggregate subgraph, and is also the basic unit for operator compilation in MindSporeAKG.
[0070] 8. Pass optimization
[0071] Pass is an important part of the compilation framework. Pass analyzes and modifies (transforms) the intermediate representation (IR); similar to pipeline operations, each pass performs a specific optimization task.
[0072] 9. MindsporeIR (MindIR)
[0073] MindIR is an intermediate representation (IR) used by the Mindspore framework.
[0074] 10. Tuning: Auto-tuning technology
[0075] It is an automatic tuning technology within the AI compilation framework. Specifically, it involves determining how tasks are distributed across cores, the order in which tensors are accessed, and how memory is allocated across storage levels. These can be represented by a series of parameters, and a specific combination of these parameter values forms a configuration (or setting). Determining this parameter combination for optimal performance is called tuning, and automating this process is called auto-tuning.
[0076] 11. Three Address Code
[0077] Three-address code is a commonly used intermediate language that compilers use to improve code conversion efficiency. Each three-address code instruction can be broken down into a four-tuple: (operator, operand 1, operand 2, result). Because each statement contains three variables, meaning each instruction has at most three operands, it is called three-address code.
[0078] 12. Reshape operation
[0079] Reshape can be understood as adjusting the shape of a tensor, but the number and relative position of the data remain unchanged. For example: a tensor with one row and two columns:
[0080] [0.2,
[0081] 0.1]
[0082] Its shape can be represented as (1, 2), where the data is 0.2, 0.1. After reshaping it to (2, 1), it will become a tensor with two rows and one column:
[0083] [0.2, 0.1].
[0084] 13. Other explanations of abstract terms
[0085] Iteration space abstraction: An N-fold nested loop is abstracted into an N-dimensional iteration space. The linear constraints on the upper and lower loop bounds of each layer constitute the constraint set of the entire iteration space. This N-dimensional iteration space corresponds to the set of values of each loop index variable in the N-fold loop. Only N-dimensional vectors that satisfy the constraints are considered valid iteration vectors.
[0086] Statement instance abstraction: Each static statement in an N-nested loop and its corresponding loop iteration vector are abstracted into a dynamic statement instance. A dynamically accessed instance of a static statement S can be written as S(i), where i represents an iteration vector of an N-nested loop.
[0087] Dataspace abstraction: Abstract an M-dimensional array into an M-dimensional dataspace. The dimensions of the dataspace can be different from the iteration space. Any iteration vector in the iteration space can be mapped to a specific location in the dataspace through data access relations.
[0088] Based on the above abstract definition, a tensor is an M-dimensional data and is abstracted into a data space. (A global tensor can be understood as a tensor that is an external input to the program and has a global scope. This corresponds to an intermediate tensor, which is generated by certain statements in the program and has a non-global scope. It can be accessed only after the statement that generated it is executed.) Statement instances access tensors by iterating over vectors. Other statements are just a verbal expression, meaning that "dependencies exist between different statements."
[0089] Dependency abstraction: Access to data space can be divided into read access and write access. The dependency between dynamic statement instances is the read and write order constraint on the data space in the program.
[0090] Dependencies are generated between different dynamic statement instances, so each dependency determines the fusion level between the different dynamic statements that generate the dependency.
[0091] The following describes the system architecture provided by the embodiments of this application. This system architecture is applied to the core feature of graph-computing fusion (GraphKernel) in the MindSpore artificial intelligence open source computing framework.
[0092] 2 , an embodiment of the present invention provides a system architecture, which includes MindSpore 201 and an Auto Kernel Generator (AKG) 202 .
[0093] MindSpore201, a next-generation deep learning framework, is based on industry best practices. It optimally matches the computing power of Ascend processors and supports flexible deployment across all scenarios, from devices to edge devices to the cloud. It pioneers a new AI programming paradigm and lowers the barrier to entry for AI development. As a new deep learning computing framework, MindSpore aims to achieve ease of development, efficient execution, and full-scenario coverage. To achieve this, MindSpore201 employs an automatic differentiation (AD) mechanism based on source code transformation (SCT), which allows complex combinations to be represented using control flow. Functions are converted into intermediate representations (IRs), which construct a computational graph that can be parsed and executed on different devices. Before execution, various hardware and software co-optimization techniques are applied to the computational graph to improve performance and efficiency in diverse scenarios, including device, edge, and cloud. MindSpore201 supports dynamic graphs, making it easier to verify the running mode. Thanks to the automatic differentiation mechanism based on source code transformation, switching between dynamic and static graph modes is simple. To effectively train large models on large datasets, MindSpore201 supports data parallelism, model parallelism, and hybrid parallel training through advanced manual configuration strategies, offering strong flexibility. Furthermore, MindSpore201 also features "auto-parallelism," which efficiently searches through a vast policy space to find a fast parallel strategy.
[0094] AKG202 optimizes operators in deep neural networks and provides automatic operator fusion under specific conditions. Working in conjunction with the graph-computation fusion capabilities of MindSpore201, AKG202 improves network performance on different hardware backends. AKG202 implements automatic scheduling optimization through polyhedral. In addition to the DSL operators provided by TVM, AKG input also supports subgraphs from graph-computation fusion and custom Python operators provided by MindSpore201. After a series of normalization passes, AKG converts the HalideIR into a schedule tree in the poly module. Automatic scheduling optimization, automatic segmentation, and memory migration are performed on the schedule tree. The tree is then returned to HalideIR for back-end instruction generation and back-end optimization. Two segmentation strategies are provided. For training scenarios, autotiling is used to achieve relatively optimal segmentation in a short time. For extreme performance optimization, tuning is provided, using evolutionary algorithms and cost models to find the optimal segmentation within the poly-assisted segmentation space.
[0095] At the layer level, MindSpore inputs the computational graph in the form of MindIR to the graph-computation fusion process. Graph-computation fusion then performs four optimization steps (also known as the frontend phase). During the first three optimization steps, many composite operators (such as softmax) are expanded into combinations of basic operators (such as add) based on mathematical formulas. These adjacent basic operators are then aggregated to form larger aggregated subgraphs, which are then subjected to common layer-specific optimizations (such as common subexpression elimination). The final phase, kernel partitioning, breaks up the aggregated subgraphs into smaller fused subgraphs, which are then passed to MindSporeAKG for code compilation and generation (also known as the backend phase). Subgraph segmentation is the fourth part of the frontend process, kernel partitioning.
[0096] However, given the current situation in which layer and operator layers are optimized separately in the industry's AI compilation framework, subgraph segmentation, a technology that connects the layer and operator layers, currently only uses layer information and manually sets some rules to predict operator layer optimization. This rule design is relatively conservative and cannot fully enable operator layer optimization, resulting in performance loss. In addition, such rules cannot be well reused and migrated across different hardware backends and different operator compilers.
[0097] To address the aforementioned technical issues, the computational graph segmentation method provided in the embodiments of this application utilizes polyhedron modeling to obtain multiple statements, then splits the computational graph based on the types of dependencies between the multiple statements. By considering the dependencies between the multiple statements obtained through polyhedron modeling during the computational graph segmentation process, the computational graph segmentation can be made more reasonable. Furthermore, information from the operator compiler is effectively used to assist in subgraph segmentation of the layer, achieving a balance between generalization and performance.
[0098] Based on this, an embodiment of the present application also provides another system architecture, namely, through the polygon-based kernel partitioning (PolyBased Kernel Paritition) in Figure 2, the replacement graph calculation is integrated with the original optimized step 4 (i.e., Kernel Partition) to perform sub-graph segmentation optimization.
[0099] PolyBased Kernel Partition consists of two parts: PolyAnalyzer, which uses a polyhedron model to perform dependency analysis on the aggregated computation graph and outputs an operator fusion analysis and a heuristic subgraph partitioning scheme; and GraphTuner, which performs online tuning of the heuristic subgraph partitioning scheme, further optimizing it for improved performance. Both parts require the MindsporeAKG operator compiler. The results of the heuristic subgraph partitioning scheme and the online tuning are passed back to the PolyBased Kernel Partition module, where they are processed into a subgraph partitioning format recognizable by graph-computation fusion and applied to MindIR for subgraph partitioning. The subsequent Mindspore compilation process proceeds as usual.
[0100] The following describes the computational graph splitting method provided in an embodiment of the present application. The method can be executed by a data processing device or by a component of the data processing device (such as a processor, chip, or chip system, etc.).
[0101] Please refer to Figure 3, which is a flowchart of a computation graph splitting method provided in an embodiment of the present application. The method may include steps 301 to 304. Steps 301 to 304 are described in detail below.
[0102] Step 301: Obtain a computation graph.
[0103] In the embodiment of the present application, there are multiple ways for the data processing device to obtain the calculation graph. The calculation graph can be obtained through the first three stages of MindSpore in the system architecture shown in Figure 2 above, or by receiving it from other devices, or by selecting it from a database, etc. The specific methods are not limited here.
[0104] The description of the computation graph can refer to the explanation of the related terms mentioned above, which will not be repeated here. In addition, the computation graph can have multiple formats (or it can be understood that the computation graph has multiple representations), which can be a MindIR computation graph or a Json type computation graph, etc., which are not limited here. MindIR is the format for representing computation graphs in the MindSpore compilation framework.
[0105] Exemplarily, the computation graph is a MindIR computation graph, and the data processing device fuses different computation graphs through the graph-computation fusion feature to generate a MindIR aggregate subgraph.
[0106] For example, Figure 4 is an example of a computational graph, in which black squares (e.g., input, input_0, input_1, input_2) are used to represent the input data of the program (i.e., the running code of the computational graph, or the program obtained after the computational graph is compiled), which can also be understood as a global tensor. White squares (e.g., 0.0001952648162841797, 1.0013580322265625e-05) are used to represent the constant input data of the program. White ellipses (e.g., Reshape, Cast, ReduceSum, Mul, Sub, Add, Rsqrt) are used to represent computational nodes. Black ellipses (e.g., out0, out1, out2, out3, output) are used to represent output tensors, arrows are used to indicate the flow of data, and the square brackets next to the arrows and the numbers inside are used to indicate the dimensions of the data. For example, if input_0 is connected to three Reshape operations, the data in the 0th dimension of input_0 will be used as the input of the Reshape operation. In the right branch, the data in the 0th dimension of the output of the Reshape operation will be used as the input of the Cast operation, and so on. Figure 4 will be used as an example in the following steps, unless otherwise specified in the subsequent steps that this example does not explain the process of the current step.
[0107] Step 302: Perform polyhedron modeling on the computation graph to obtain multiple statements.
[0108] After the data processing device obtains the computation graph, it performs polyhedral modeling on the computation graph to obtain multiple statements. The polyhedral modeling process can be understood as a process of constructing a mathematical model. For example, polyhedral modeling can describe a program using iteration space, statement instances, access relationships, dependencies, and scheduling, and implement a series of parallelism and locality-related optimizations based on scheduling transformations. Operators in the computation graph are typically represented as tensor calculations within nested loops during their actual execution at the underlying layer. The task of polyhedral modeling is to convert these operators into code suitable for efficient execution on the underlying architecture through loop transformations.
[0109] Among them, the description of the statements in the above multiple statements can refer to the explanations in the aforementioned related terms, which will not be repeated here.
[0110] In one possible implementation, the computation graph is a MindIR computation graph, and the data processing device first converts the computation graph into a computation script, and then uses the computation script to perform polyhedron modeling to obtain multiple statements.
[0111] For example, the data processing device converts the MindIR computation graph into a computation script (e.g., a JSON-type computation script) using MindSpore's built-in conversion tool. The computation script is then modeled using a polyhedron model to generate multiple statements. For example, a JSON-type computation script can describe at least one of the following information for each specific operator in the computation graph: input, output, operation, and the execution order between operators.
[0112] It is understandable that in order to achieve a finer-grained splitting of the subsequent computation graph, each operator in the computation graph will be abstractly modeled using a polyhedron model, and a Poly statement (also called statement instance abstraction) will be generated during the modeling process.
[0113] For example, continuing with the example of Figure 4 above, the calculation graph of Figure 4 is first converted into a calculation script of the Json type as shown in Figure 5. Among them, lines 41 to 82 of Figure 5 describe the Reshape operation, which includes the input (input_desc), output (output_desc), operation (name), attributes related to the operation (attr), and some memory-related content (ptr_address). All operations will be arranged in a legal topological order and placed in the op_desc list on line 40. Then, AKG analysis is performed on the calculation script to generate HalideIR. Specifically, the calculation is first translated into IR in the form of three-address code according to the topological order in Json, for example: "output0_0(12288,5120)=Reshape(input_0(12,1024,5120)):float16:PI". This three-address IR is then translated into HalideIR, the intermediate language used by AKG. This step directly calls the API provided by TVM. The computation in the three-address IR can be described as a loop operation, including memory allocation, loop declaration, and the specific operations within each loop. Notably, the destination tensor naming logic can adopt the format "T_computation_source tensor name." For example, if the destination tensor name is "T_reshapeinput_0," its operation is reshape, and the source tensor name is input_0. Because TVM is used to translate the three-address IR into HalideIR, TVM also follows a topological order when translating compute nodes into IR. This topological order may not match the JSON representation. Therefore, before translating the IR, the name of each tensor in HalideIR and its topological order in JSON are recorded and annotated in the HalideIR. This HalideIR is then parsed into statements by the polyhedron model. For example, the mapping relationship between statement S_0 in the Poly log and HalideIR is "S_0: T_reshapeinput_0(ax0, ax1) = input_0(floordiv(ax0, 1024), floormod(ax0, 1024), ax1)". Based on this mapping relationship, we can obtain the mapping of Json topological sequence -> HalideIR statement -> Poly statement, as shown below:
[0114] Step 303: Split the computation graph based on the types of dependency relationships between the multiple statements to obtain multiple first subgraphs.
[0115] After obtaining multiple statements, the data processing device may split the computation graph based on the type of dependency relationships between the multiple statements to obtain multiple first subgraphs. The dependency relationships are used to represent the order of reads and / or writes between the multiple statements. For example, a read-write order relationship, a write-write order relationship, etc.
[0116] This step can also be understood as first determining the type of dependency relationship between multiple statements, and then splitting the computation graph based on the type of dependency relationship to obtain multiple first subgraphs.
[0117] Optionally, statement dependencies include computation dependencies and / or data dependencies. Computation dependencies represent the relationship between memory address units referenced by multiple statements. Data dependencies represent the relationship between execution paths of multiple statements at runtime.
[0118] For example, if the statement dependency includes computation dependency and data dependency, and the multiple statements include S1 statement and S2 statement, the dependency of S2 statement on S1 can be understood as follows:
[0119] 1. Computational dependencies: At runtime, there may be an execution path from statement S1 to statement S2.
[0120] 2. Data dependency: Statement S2 and statement S1 reference the same memory address unit, and at least one of the statements S2 and S1 has a write reference type.
[0121] Optionally, taking into account the scheduling characteristics of the polyhedron model, the types of dependencies in the embodiments of the present application include at least one of the following: a first type, a second type, a third type, and a fourth type, etc. Among them, the first type is used to indicate that the access order between the statements is the same, and the number of instances of write operations and read operations between the statements is the same. The second type is used to indicate that the access order between the statements is different, and the number of instances of write operations and read operations between the statements is the same. The third type is used to indicate that the number of write operations of each statement is greater than or equal to the number of read operations. The fourth type is used to indicate that the number of write operations of each statement is less than the number of read operations. The fifth type is used to indicate an injective in a mathematical mapping (i.e., the mapping relationship between the two statements is one-to-one). The sixth type is used to indicate a non-injective in a mathematical mapping (i.e., the mapping relationship between the two statements is not one-to-one, for example, many-to-one, one-to-many, many-to-many, etc.).
[0122] Types 1 to 4 can be understood as computational dependency types, while types 5 and 6 are data dependency types. For ease of understanding, type 1 can be referred to as OneToOne, type 2 as OneToDiffOne, type 3 as OneToMany, type 4 as ManyToOne, type 5 as Injective, and type 6 as NonInjective.
[0123] Optionally, after determining the type of dependency relationship between multiple statements, it can be saved in a dictionary format similar to "the dependency relationship between statement A and statement B is OneToOne".
[0124] For example, continuing with the above example, since there may be dependencies between statements and global tensors and other statements, the polyhedron model can describe these dependencies, for example: "{S_4[ax0, ax1]->input_0[arg0, arg1, arg2=ax1]:(-ax0+arg1)mod1024=0and0<=ax0<=12287and0<=ax1<=5119and0<=arg0<=11and-1023+ax0<=1024arg0<=ax0and0<=arg<=1023}". Among them, S_4[ax0, ax1]->input_0[arg0, arg1, arg2=ax1] represents the following three points:
[0125] 1. S_4 has a dependency relationship with input_0. The dependency relationship is that ax1 of S_4 and arg2 of input_0 are in a one-to-one mapping relationship (refer to the relationship between ax1 of the tensor on the left and ax1 of the tensor on the right in line 1085);
[0126] 2. The ax0 of S_4 is related to the arg0 and arg1 of input_0. The calculation logic is (-ax0 + arg1) mod 1024 = 0, and the range of ax0 is [0, 12287], and the range of arg1 is [0, 1023]. That is, when arg1 takes any value in the range of [0, 1023], ax0 must be equal to arg1, otherwise the equality does not hold (refer to the relationship between ax0 of the tensor on the left and floormod (ax0, 1024) of the tensor on the right in line 1085);
[0127] 3. Similarly, when arg0 takes any value in the range [0, 11], ax0 must be equal to 1024*arg0 (refer to the relationship between ax0 of the left tensor and floordiv(ax0, 1024) of the right tensor in line 1085).
[0128] It can be seen that based on the dependency relationship of the polyhedron model, the calculation logic on HalideIR can be completely restored.
[0129] For example, the dependency relationship between statements S_4 and S_5 is: "{S_4[ax0, ax1]->S_4[ax0'=ax0, ax1'=ax1]: 0<=ax0<=12287and 0<=ax1<=5119}". Among them, "S_4[ax0, ax1]->S_5[ax0'=ax0, ax1'=ax1] means that S_5 depends on S_4, ax0' of S_5 is equal to ax0 of S_4, and ax1' of S_5 is equal to ax1 of S_4".
[0130] In this step, the dependencies analyzed by poly can be matched with the dependents in these dependencies to generate the following format:
[0131] The following example describes how to determine the type of dependency between statements.
[0132] For any statement's dependencies between statements and global tensors, the steps to obtain dependency classification are as follows:
[0133] Step 1: Check whether the dimensions of the dependency are consistent, that is, whether the number of parameters in the square brackets after the statement / global tensor name is consistent:
[0134] For example, dependency 1 indicates that the dimensions of S_4 and S_5 are consistent. For example, dependency 1 is "S_4[ax0, ax1]->S_5[ax0'=ax0, ax1'=ax1]: 0<=ax0<=12287 and 0<=ax1<=5119".
[0135] For another example, dependency 2 indicates that the dimensions of S_0 and input_0 are inconsistent. For example, dependency 2 is "S_0[ax0, ax1]->input_0[arg0, arg1, arg2=ax1]: (-ax0+arg1) mod 1024=0 and 0<=ax0<=12287 and 0<=ax1<=5119 and 0<=arg0<=11 and -1023+ax0<=1024arg0<=ax0 and 0<=arg1<=1023".
[0136] If the dimensions are consistent, proceed to step 2. Otherwise, if it is a dependency between a statement and a global tensor, directly determine that the dependency is of type NonInjective (i.e., type 6). If it is a dependency between statements, proceed to step 3.
[0137] Step 2: Check whether access to the same data block in the dependency is continuous. For dependencies between statements, check whether the order of data access by the dependent axes is consistent.
[0138] For example, an example of data access being continuous is “S_0[ax0, ax1]->S_1[ax0′=ax0, k1=ax1]: 0<=ax0<=1 and 0<=ax1<=2”.
[0139] Dependencies between statements with sequential access (consistent order) are classified as OneToOne dependencies (Type 1). Figure 6 shows a diagram of the Type 1 iteration space: their iteration spaces completely overlap, meaning they have the same number of computation instances in their iteration spaces and, without scheduling, their access order is also consistent. Dependencies between statements with sequential access (consistent order) and global tensors are classified as Injective dependencies (Type 5).
[0140] For another example, the access is non-continuous (which can be understood as a transpose) as "S_0[ax0, ax1]->S_1[ax0'=ax1, ax1'=ax0]: 0<=ax0<=1 and 0<=ax1<=2".
[0141] Dependencies between statements with non-contiguous access (inconsistent order) are classified as the OneToDiffOne dependency type (i.e., type 2). Figure 7 shows a diagram of the second type of iteration space: two statements have the same number of computational instances in their iteration space, but the access order is inconsistent. Dependencies between statements with non-contiguous access (inconsistent order) and global tensors are classified as the NonInjective dependency type (i.e., type 6).
[0142] Step 3: Check the number of write instances and read instances in the dependency. Note: An instance represents an operation on a specific location. A read instance reads data from a location in a tensor, while a write instance writes data to a location in a tensor.
[0143] If the number of instances of the write operation is the same as the number of instances of the read operation, then it is determined to be the OneToDiffOne type (i.e., the second type). For example, the number of read and write accesses in the dependency relationship in Figure 8 is the same, but the dimensions are inconsistent (which can be understood as a reshape). Figure 8 corresponds to: "S_0[ax0]->S_1[arg0, arg1]: (-ax0+arg1)mod 1024=0 and 0<=ax0<=12287 and 0<=arg0<=11 and -1023+ax0<=1024arg0<=ax0 and 0<=arg1<=1023".
[0144] If the write instance is greater than the read instance, that is, one read instance corresponds to multiple write instances, then the operation is determined to be of the OneToMany type (i.e., the third type). For example, a diagram of the third type iteration space is shown in Figure 9: the dependency S_1 read instance is on ax0, and the write instance is on ax0 and ax1 (which can be understood as a broadcast of the ax1 axis). Figure 9 corresponds to: "S_0[ax0]->S_1[ax0'=ax0,ax1]: 0<=ax0<=1 and 0<=ax1<=2".
[0145] If the number of write instances is smaller than the number of read instances, meaning that multiple read instances correspond to one write instance, then the relationship is considered ManyToOne (i.e., the fourth type). For example, the following dependency is a special self-dependency (one that depends on itself). A schematic diagram of the fourth type of iteration space is shown in Figure 10 (which can be understood as a reduce operation). Figure 10 corresponds to: "S_0[ax0, k1]->S_0[ax0'=ax0, k1'=1+k1]: 0<=ax0<=1 and 0<=k1<=2".
[0146] According to the above steps, you can update the analysis results DataDependenciesType and ComputeDependenciesType of each statement.
[0147] For example, since S_0 is the first statement, it is the only one that depends on the result of the global tensor. To include the analysis of the dependencies between the above statements, we use S_5 as an example:
[0148] The above describes the type of dependency determination. The following describes the application of this type. That is, after the data processing device determines the type of dependency between multiple statements, it can split the computational graph based on the type of dependency to obtain multiple first subgraphs.
[0149] Optionally, the data processing device first determines a fusion level between statements of corresponding types based on a mapping relationship, where the mapping relationship describes the correspondence between the type and the fusion level, and the fusion level describes the fusibility of each statement. The computation graph is then split based on the fusion level to obtain multiple first subgraphs.
[0150] Optionally, in the embodiment of the present application, the fusion level includes at least one of the following: the first level, the second level, the third level, and the fourth level. Among them, the first level is used to indicate that each statement can be fused front and back. The second level is used to indicate that each statement must not be fused back. The third level is used to indicate that each statement must not be fused front and back. The fourth level is used to indicate that each statement may be fused front and back. For ease of understanding, the first level is called front and back fusion (Fusible), the second level is called front fusion (Infusible), the third level is called non-fusion (StandAlone), and the fourth level is called unknown fusion (MayFuse).
[0151] Exemplarily, the mapping relationship between the dependency types and fusion levels includes: the correspondence between the first level and the first type, the correspondence between the first level and the fifth type, the correspondence between the second level and the fourth type, the correspondence between the third level and the first combination type, the correspondence between the fourth level and the second type, the correspondence between the fourth level and the third type, and the correspondence between the fourth level and the sixth type; the first combination type includes the combination of the second type and the fourth type. An example of the mapping relationship is shown in Table 1:
[0152] Table 1
[0153] Therefore, according to the mapping relationship in Table 1, we can first check the dependency between statements. If it does not exist, we will check the dependency between the statement and the global tensor to update the FuseDecisionMap of the analysis result of each statement. The specific example is as follows:
[0154] This step allows you to obtain analysis results for all statements. The analysis results can be represented in a dictionary format, for example, they can be saved as a file in JSON format. Subsequent steps will read the JSON file on the Python side to obtain the polyhedron model analysis results before proceeding.
[0155] The above describes the types of dependency relationships and fusion levels. The following describes in detail how to perform subgraph segmentation with reference to FIG11 . The process includes steps 1101 to 1104 .
[0156] Step 1101: Initialize subgraph segmentation.
[0157] Perform polyhedral scheduling on the aggregated subgraph, and initialize the statements contained in each filter (which can be understood as a container similar to a list) after scheduling into a subgraph. The statements after polyhedral scheduling will be given in the form of a scheduling tree.
[0158] For example, the scheduling tree after scheduling is as follows:
[0159] The set and sequence nodes represent whether the filter nodes defined below them are executed out of order (set) or in order (sequence). For example, if the second line is a set, and its child nodes are the sequence node on the third line and the filter node containing the S_0 statement on line 33, this means that all statements on lines 3-32 have no dependencies on the S_0 statement and can be executed out of order (it doesn't matter which one executes first). For example, the sequence node on line 9 indicates that the filter nodes on lines 10 through 13 must be executed sequentially, so statements S_4 must be executed before S_5, followed by S_6 and S_7.
[0160] Based on this scheduling result, the statements in the filter nodes (such as the filter nodes in lines 10-13, 17-23, and 29-32) contained under all sequence nodes that only contain filter child nodes (such as the sequence nodes in lines 9, 16, and 28) can be initialized into a subgraph and given in list form. As a result, the following three subgraphs can be obtained (also called the initialization subgraph stage 1):
[0161] 1.[S_4, S_5, S_6, S_7]
[0162] 2.[S_13, S_3, S_9, S_10, S_11, S_12, S_14]
[0163] 3.[S_18, S_19, S_20, S_21]".
[0164] Next, initialize the statements contained in other childless filter nodes into a subgraph and arrange them in the order of the scheduling tree. If the order cannot be determined for set nodes, sort them from smallest to largest by statement index (for example, the index of statement S_0 is 0). This results in the following subgraph (also called the initialization subgraph stage 2). The row numbers before these subgraphs can also be treated as the subgraph indices, and they are in order.
[0165] “1.[S_0]
[0166] 2.[S_1]
[0167] 3.[S_2]
[0168] 4.[S_4, S_5, S_6, S_7]
[0169] 5.[S_8]
[0170] 6.[S_13, S_3, S_9, S_10, S_11, S_12, S_14]
[0171] 7.[S_15]
[0172] 8.[S_16]
[0173] 9.[S_17]
[0174] 10.[S_18, S_19, S_20, S_21].
[0175] Finally, based on the relationship between the statements and the Json computation node topology derived previously, statement S_x is mapped back to the computation node Op_x. Since some statements cannot be mapped back (for example, the reduce initialization statement, which was newly created during the polyhedron modeling phase), in this case, the newly added statement can be identified based on the statement and computation mapping table, and its type can be determined based on its dependency relationship. If the statement is dependent on by the reduce statement, it is determined to be a reduce initialization statement and labeled Op_x_red_init, where x is the Json computation topology of the reduce statement. This results in the initial stage of the subgraph partitioning scheme (also known as initialization subgraph stage 3) as shown below:
[0176] Step 1102: Sub-image fusion.
[0177] Since many subgraphs with only one node are generated in the previous step 1101, and they come from sequence nodes, that is, they have dependencies, this step 1102 will try to fuse the subgraphs containing only one statement in descending order of subgraph index based on the fusionability between statements.
[0178] The fusion strategy is to merge two statements into the subgraph containing the closest dependent statement when the dependency type is Fusible, and place it before the dependent statement. For example, in the above example, we first try to fuse Op15 and its dependent statements. Since Op16 depends on Op15 and the dependency type is Fusible, the subgraph containing Op15 and Op16 is fused, and Op15 is placed before Op16. The heuristic subgraph solution becomes as follows (also known as the initialization subgraph stage 4):
[0179] Then comes Op14, Op13, and so on. The final integrated heuristic subgraph solution is as follows (also known as the parallelism-first heuristic subgraph solution):
[0180] After fusing the single-node subgraph generated by the sequence node, an attempt is made to fuse the single-node subgraph generated by the set node. In this example, Op0 corresponding to S_0 is a single-node subgraph generated by the set node, so Op0 will attempt to fuse with another unique subgraph that has a common global tensor dependency (if any). This fusion can also reduce data movement. Since Op0, Op3, and Op4 all depend on the global tensor input_0, and Op3 and Op4 are in different subgraphs, it is impossible to determine the fusion location. Therefore, no fusion will be performed in this step, only annotation will be performed. During the online automatic tuning phase of step 304 (if any), a tuning attempt will be made to see which subgraph has better fusion performance.
[0181] Step 1103: Rearrange the topology within the subgraph.
[0182] Adjust the order of nodes within the subgraph based on fusibility. While ensuring correctness, rearrange the topological order of statements within the subgraph according to the priority of "Fusible > MayFuse > Infusible > StandAlone." Statement types with higher priorities have lower topological orders within the subgraph. This step prepares the next step for generating as few subgraphs as possible. For example, consider three computations: Op0, Op1, and Op2. Op0 depends on input, Op1 and Op2 both depend on Op0, and Op0 is fusible with the input, Op1 is infusible with Op0, and Op2 is fusible with Op0. Then, a subgraph consisting of these three computation nodes can be represented as "[Op0(Fusible), Op1(Infusible), Op2(Fusible)]."
[0183] Since Op1 and Op2 do not depend on each other, Op0, Op2, Op1 is also a legal topological order; since the fusible nature of Op2 is greater than the infusible nature in the reordering priority, after this stage, the subgraph will be reordered to "[Op0(Fusible), Op2(Fusible), Op1(Infusible)]".
[0184] Since the above example of the heuristic subgraph solution with parallelism priority (ie, the example in step 1102 ) does not involve this situation, the subgraph partitioning solution remains unchanged.
[0185] Step 1104: subgraph segmentation.
[0186] According to the aforementioned definition, StandAlone nodes do not allow for front-end fusion, and Infusible nodes do not allow for back-end fusion. Therefore, this strategy can be followed to split each rearranged subgraph. Note: Only Infusible and StandAlone nodes are split, and the Fusible and MayFuse nodes are fused. This can reduce the number of kernels. For example, for the example "[Op0(Fusible), Op1(Infusible), Op2(Fusible)]" in step 1103 above, splitting will generate the following two subgraphs:
[0187] 1.Micro graph 0":[Op0(Fusible),Op1(Infusible)],
[0188] 2."Micro graph 1":[Op2(Fusible),].
[0189] However, if the example "[Op0(Fusible), Op2(Fusible), Op1(Infusible)]" in step 1103 is split, there is no change (the Infusible node is at the end of the graph and does not need to be split).
[0190] For example, a graph containing StandAlone (e.g. "Micro graph 0": [Op0(Fusible), Op1(StandAlone), Op2(Fusible)]") will be split into the following three graphs:
[0191] 1."Micro graph 0":[Op0(Fusible),],
[0192] 2."Micro graph 1":[Op1(StandAlone),],
[0193] 3."Micro graph 2":[Op2(Fusible),].
[0194] Since all the infusible nodes Op6 and Op12 in the last example of step 1102 are exactly at the end of the subgraph, the subgraph partitioning scheme remains unchanged.
[0195] After these four steps, the obtained subgraph segmentation scheme (ie, the last example of step 1102 ) is the final heuristic subgraph segmentation scheme.
[0196] Optionally, after obtaining the subgraph segmentation scheme, the obtained segmentation scheme can be converted into the format used in the original graph-algorithm fusion process in Mindspore and returned to Mindspore graph-algorithm fusion. For example, the last example in step 1102 will be converted into the following format (i.e., the subgraph segmentation result representation in graph-algorithm fusion format):
[0197] Among them, group_num is used to indicate the number of graphs that were finally cut out. Here, four graphs were cut out, numbered 0, 1, 2, and 3. split_result represents the splitting plan, which is a list with the same length as the number of calculation nodes in the Json. The value at each position represents the subgraph number to which the calculation of the corresponding topological order belongs. For example, the 0th position is 3, which means that the 0th node (Op0) belongs to the 3rd subgraph.
[0198] Step 304: Optimize the multiple first subgraphs to obtain multiple second subgraphs. This step is optional.
[0199] Optionally, after acquiring the multiple first subgraphs, the data processing device may further optimize the multiple first subgraphs to obtain the multiple second subgraphs.
[0200] This step can also be understood as further splitting the first subgraph, and the number of the multiple second subgraphs is greater than or equal to the number of the multiple first subgraphs. The multiple second subgraphs are used to improve the processing performance of the data processing device.
[0201] When there are tuning conditions and requirements, each MicroGraph in the above example (i.e., the last example in step 1102) can be treated as an independent Task, labeled Task0, Task1, Task2, and Task3. For each Task, online automatic tuning is performed. The specific process is shown in Figure 12. This process includes steps 1201 to 1204, which are described below:
[0202] Step 1201, generate space.
[0203] Optionally, this step may be performed by the space generator (SpaceGenerator) in FIG2 .
[0204] Specifically, a search algorithm (such as Depth-First-Search (DFS)) can be used to traverse each Op in the Task. During this process, a queue of the current subgraph and a queue of the total subgraph are maintained, and the compatibility of the current Op with its previous Ops is checked. The specific situation is as follows:
[0205] The first case: If there is a previous Op that is in a StandAlone relationship with the current Op, this operation means that the Op can never be merged before (there can be no other Op before it) and can never be merged after (there can be no other Op after it). In this case, the entire current subgraph queue is placed into the total subgraph queue, and the current subgraph queue is cleared. Then, the current Op is taken as a complete subgraph queue, placed into the total subgraph queue, and the next Op is traversed.
[0206] The second case: If there is a previous Op that is in an infusible relationship with the current Op, this operation means that the Op can always be fused forward (there can be other Ops before it) and can never be fused backward (there can be no other Ops after it). In this case, after adding the current Op to the current subgraph queue, the entire current subgraph queue is placed in the total subgraph queue, and after clearing the current subgraph queue, traversing the next Op continues.
[0207] The third case: If there is a previous Op that is in a fusible relationship with the current Op, this operation means that the Op can always be fused forward (there can be other Ops before it) and can always be fused backward (there can be other Ops after it). In this case, the current Op is added to the current subgraph queue and the next Op is traversed.
[0208] The fourth case: If none of the above three conditions are met, there may be two cases:
[0209] 1. The Op does not depend on any previous Op, it only depends on the global tensor, so it is processed like a Fusible.
[0210] 2. If there is a previous Op in a MayFuse relationship with the current Op, this means that the Op cannot always be front-fused (it is uncertain whether there are other Ops before it) and cannot always be back-fused (it is uncertain whether there are other Ops after it). In this case, three states are saved for the current Op:
[0211] State 1: Allows pre- and post-fusion, then the Op is processed like a Fusible.
[0212] State 2: Only pre-fusion is allowed, and post-fusion is not allowed. In this case, the Op is handled like Infusible.
[0213] State 3: Pre- and post-fusion are not allowed: Then the Op is processed like StandAlone.
[0214] For each of these three states, continue traversing the next Op. Once the current subgraph is traversed, the queue of all non-repeated subgraphs is saved, creating a candidate subgraph partitioning plan (hereinafter referred to as a candidate plan). This process (using an example of DFS generating a space) can be seen below (for simplicity, another simple graph is used, unrelated to the previous example):
[0215] Step 1202, search.
[0216] Optionally, this step may be performed by the GraphTuner in FIG2 .
[0217] Specifically, for each candidate plan, a brute force search (traversing each plan) is selected if time permits; a random search (randomly selecting N candidate plans) is selected if time does not permit.
[0218] Step 1203: Code generation and evaluation.
[0219] Optionally, this step can be performed by the GraphRunner in FIG2 .
[0220] Specifically, for each candidate plan, a corresponding Json calculation graph that can be run by AKG is generated based on the Ops contained therein. Then, the Json calculation graph is input into AKG for normal operator compilation and running process. The running time of the calculation graph is obtained and recorded.
[0221] For example, for a MicroGraph like [7, 3, 8, 9, 10, 11, 12], the fastest split is [[7], [3, 8, 9, 10, 11, 12]]. This means that node 7 is treated as a separate subgraph, and the other nodes are treated as a subgraph. Their running times are 13.131 microseconds (us) and 210.553 us, respectively. The total running time of this split is 13.131 + 210.553 = 223.684. For example, the running time rankings of the various plans in Task 2 are as follows:
[0222] Finally, for each task, the candidate plan with the shortest total time is selected to form the final subgraph partitioning plan. If the total time is the same for both tasks, the one with the smallest number of subgraphs is selected. If the number of subgraphs is the same, a random one is selected.
[0223] For example, “[[0]]”:”[334.202,[334.202]]” is the subgraph segmentation scheme selected by Task0, “[[4,5,6]]”:”[99.485,[99.485]]” is the subgraph segmentation scheme selected by Task1, Figure 13 is the subgraph segmentation scheme selected by Task2, and Figure 14 is the subgraph segmentation scheme selected by Task3.
[0224] Finally, the above example aggregate subgraph segmentation scheme is as follows (hereinafter referred to as stage 1):
[0225] [0], [4, 5, 6], [7], [3, 8, 9, 10, 11, 12], [13, 14, 15, 1, 2, 16], [17, 18, 19].
[0226] Step 1204, MicroGraph fusion.
[0227] Optionally, this step may be performed by the GraphTuner in FIG2 .
[0228] Specifically, first, based on the code generation capability of MindsporeAKG, the subgraph containing Infusible can also be compiled separately from its consumer subgraph using a segmented compilation method. The compiled IR is then spliced to achieve the fusion effect (this function is called BufferStitch in the graph fusion feature). Therefore, based on this function, the graph [4, 5, 6] in the subgraph segmentation scheme in stage 1 will try to merge with [7], and [3, 8, 9, 10, 11, 12] and [13, 14, 15, 1, 2, 16] will try to merge. The trial process is: obtain the running time after merging, compare it with the running time without merging, and save the better one. If the time is the same, the merging solution (with fewer subgraphs) is adopted. For example, Table 2 below is the data obtained during the attempted merging of [4, 5, 6] and [7]. After comparison, the solution [4, 5, 6, 7] is finally selected:
[0229] Table 2
[0230] Furthermore, based on MindsporeAKG's polyhedron scheduling capabilities, subgraphs that have no dependencies on other nodes will attempt to merge with subgraphs containing statements that depend on the same global tensor (see step 4 in section 3.6.1.1.4). Therefore, based on this feature, the graph [0] will attempt to merge with [4, 5, 6, 7] or [3, 8, 9, 10, 11, 12]. Similar to the process in "First" above, the splitting scheme with the best performance and the least number of subgraphs will be selected. Therefore, after comparison, the scheme [0, 4, 5, 6, 7] was ultimately selected.
[0231] The final aggregate subgraph segmentation plan for the above example (which can be called stage 2) is: "[0], [4, 5, 6], [7], [3, 8, 9, 10, 11, 12], [13, 14, 15, 1, 2, 16], [17, 18, 19]".
[0232] In the embodiment of the present application, on the one hand, multiple statements are obtained by polyhedron modeling of the computation graph, and then the computation graph is split according to the type of dependency relationship between the multiple statements. In the process of splitting the computation graph, the dependency relationship between the multiple statements obtained by polyhedron modeling is considered, which can make the splitting of the computation graph more reasonable. Specifically, compared with the dependency relationship of the modeling of the prior art, which is based on the operator granularity, the dependency relationship obtained by the embodiment of the present application based on the polyhedron model is based on the loop in the operator granularity. The dependency relationship obtained in this way is more specific and has the characteristics of operator generation (operator generation is also scheduled based on this loop dependency). Therefore, it is more reasonable to use this dependency relationship to divide the computation graph and subsequent operator generation. In addition, the information of the operator compiler is well used to help the layer to perform subgraph segmentation, taking into account generalization and performance. On the other hand, compared with the existing search space explosion that cannot achieve subgraph tuning, the solution of the embodiment of the present application splits the computation graph according to the type of dependency relationship, so that multiple first subgraphs can be further optimized to obtain multiple second subgraphs. That is, tuning can be completed in a short time and reach industrial usability level. And based on the fusibility, the tuning space is pruned in large quantities, which solves the problem of the inability to perform automatic tuning in the subgraph segmentation problem in the prior art. On the other hand, compared with the AI compilation framework in the prior art, the layer and operator layer are optimized separately, and the information cannot be synchronized between the two levels. The solution of the embodiment of the present application can pass the information of the layer, such as the topological order in the calculation graph, to the operator layer. On the other hand, based on the ability of graph-calculation fusion to expand complex operators into basic operators, statement fusion is canceled through compilation options, and topological order annotation is performed during the operator translation process to IR, which solves the problem that information cannot be shared between the layer and operator layers.
[0233] In order to more intuitively see the beneficial effects of the embodiment of the present application, the results of the sub-graph splitting scheme performed by the prior art and the embodiment of the present application are analyzed below. The analysis results are shown in Table 3.
[0234] Table 3
[0235] It can be seen that compared with the prior art, the embodiment of the present application reduces the number of subgraphs by more than half and improves performance by 44%. This is due to two factors. On the one hand, based on the fusion determination of the polyhedron model, more statements are determined to be capable of operator fusion (for example, Op1 and Op2 are both fused with their consumer subgraphs in the solution of the embodiment of the present application). On the other hand, MicroGraph merging attempts are performed during online automatic tuning.
[0236] In addition, the following compares the applicability of the prior art and the methods provided in the embodiments of this application to different backends (e.g., the backends supporting polyhedron models in MindsporeAKG). Specifically, examples include Ascend, central processing units (CPUs), and graphics processing units (GPUs). Figure 15 shows a performance comparison in a GPU scenario, and Figure 16 shows a performance comparison in a CPU scenario.
[0237] Among them, the vertical coordinate of Figure 15 is 1, which is a handwriting library commonly used in the prior art, the vertical coordinate of 1.2 is the prior art, and the vertical coordinate of 1.42 is steps 301 to 303 (which can be called MS+poly split) in the method provided in the embodiment of the present application. The vertical coordinate of 1.46 is steps 301 to 304 (which can be called MS+poly split+tune) in the method provided in the embodiment of the present application. It can be seen that the method provided in the embodiment of the present application improves the performance of the entire network better than the prior art. That is, the performance of the heuristic algorithm (poly split) is better than the rule-based (rule-base split) solution:
[0238] 1. GPU large model training can achieve up to 1.42x acceleration;
[0239] 2. CPU inference scenarios can achieve up to 2.5x acceleration.
[0240] Furthermore, the method provided in the embodiment of the present application can also be combined with the tuning framework for ultimate tuning:
[0241] 1. GPU large model training can achieve up to 1.46x acceleration, and the entire network tuning time is 3 hours (industrially acceptable);
[0242] 2. CPU inference scenarios can achieve up to 2.86x acceleration, and the entire network tuning time is 15 minutes.
[0243] The above describes the computational graph splitting method in the embodiment of the present application. The following describes the data processing device in the embodiment of the present application. Please refer to Figure 17. An embodiment of the data processing device in the embodiment of the present application includes:
[0244] An acquisition unit 1701 is used to acquire a computation graph, where the computation graph is used to represent the computation logic of multiple operator nodes.
[0245] A modeling unit 1702 is configured to perform polyhedron modeling on the computation graph to obtain multiple statements;
[0246] The splitting unit 1703 is used to split the computation graph based on the type of dependency relationship between multiple statements to obtain multiple first subgraphs, where the dependency relationship is used to represent the sequential relationship of reading and / or writing between multiple statements.
[0247] Optionally, the data processing device further includes: an optimization unit 1704, configured to optimize the plurality of first subgraphs to obtain a plurality of second subgraphs.
[0248] Optionally, the type of dependency includes at least one of the following: a first type, a second type, a third type, a fourth type, a fifth type, and a sixth type; the first type is used to indicate that the access order between statements is the same, and the number of instances of write operations and read operations between statements is the same; the second type is used to indicate that the access order between statements is different, and the number of instances of write operations and read operations between statements is the same; the third type is used to indicate that the number of write operations of each statement is greater than the number of read operations; the fourth type is used to indicate that the number of write operations of each statement is less than the number of read operations; the fifth type is used to indicate that the mapping relationship between statements is one-to-one; and the sixth type is used to indicate that the mapping relationship between statements is not one-to-one.
[0249] Optionally, the splitting unit 1703 is specifically used to determine the fusion level between statements with corresponding types based on the mapping relationship, the mapping relationship is used to describe the correspondence between the type and the fusion level, and the fusion level is used to describe the fusibility of each statement; the splitting unit 1703 is specifically used to split the computational graph based on the fusion level to obtain multiple first subgraphs.
[0250] Optionally, the fusion level includes at least one of the following: first level, second level, third level and fourth level; the first level is used to indicate that each sentence can be fused front to back, the second level is used to indicate that each sentence cannot be fused back to back, the third level is used to indicate that each sentence cannot be fused front to back, and the fourth level is used to indicate that each sentence can be fused front to back.
[0251] Optionally, the mapping relationship includes: the correspondence between the first level and the first type, the correspondence between the first level and the fifth type, the correspondence between the second level and the fourth type, the correspondence between the third level and the first combination type, the correspondence between the fourth level and the second type, the correspondence between the fourth level and the third type, and the correspondence between the fourth level and the sixth type; the first combination type includes the combination of the second type and the fourth type.
[0252] In this embodiment, the operations performed by each unit in the data processing device are similar to those described in the embodiments shown in Figures 1 to 16 above, and will not be repeated here.
[0253] In this embodiment, in the embodiment of the present application, the computation graph is polyhedron modeled by the modeling unit 1702 to obtain multiple statements, and the splitting unit 1703 then splits the computation graph according to the type of dependency relationship between the multiple statements. In the process of splitting the computation graph, the dependency relationship between the multiple statements obtained by the polyhedron modeling is considered, so that the splitting of the computation graph can be made more reasonable.
[0254] Referring to FIG. 18 , this application provides a schematic diagram of another data processing device. The data processing device may include a processor 1801, a memory 1802, and a communication port 1803. The processor 1801, the memory 1802, and the communication port 1803 are interconnected via a circuit. The memory 1802 stores program instructions and data.
[0255] The memory 1802 stores program instructions and data corresponding to the steps executed by the data processing device in the corresponding implementation modes shown in Figures 1 to 16 above.
[0256] The processor 1801 is configured to execute the steps performed by the data processing device in any of the embodiments shown in FIG. 1 to FIG. 16 .
[0257] The communication port 1803 can be used to receive and send data, and to execute the steps related to acquisition, sending, and receiving in any of the embodiments shown in Figures 1 to 16 above.
[0258] In one implementation, the data processing device may include more or fewer components relative to FIG18 . This application is merely an illustrative description and is not intended to be limiting.
[0259] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0260] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0261] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in whole or in part through software, hardware, firmware, or any combination thereof.
[0262] When software is used to implement the integrated unit, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0263] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
Claims
1. A computational graph splitting method, characterized in that: The method comprises: Obtaining a computational graph, where the computational graph is used to represent computational logic of multiple operator nodes; Performing polyhedral modeling on the computation graph to obtain a plurality of statements; The computation graph is split based on the type of dependency relationship between the multiple statements to obtain multiple first subgraphs, and the dependency relationship is used to represent the sequential relationship of reading and / or writing between the multiple statements.
2. The method according to claim 1, characterized in that The method is applied to a data processing device, and the method further comprises: The plurality of first subgraphs are optimized to obtain a plurality of second subgraphs, wherein the number of the plurality of second subgraphs is greater than or equal to the number of the plurality of first subgraphs.
3. The method according to claim 1 or 2, characterized in that: The types include at least one of the following: a first type, a second type, a third type, a fourth type, a fifth type and a sixth type; the first type is used to indicate that the access order between statements is the same, and the number of instances of write operations and read operations between statements is the same; the second type is used to indicate that the access order between statements is different, and the number of instances of write operations and read operations between statements is the same; the third type is used to indicate that the number of write operations of each statement is greater than the number of read operations; the fourth type is used to indicate that the number of write operations of each statement is less than the number of read operations; the fifth type is used to indicate that the mapping relationship between statements is one-to-one; the sixth type is used to indicate that the mapping relationship between statements is not one-to-one.
4. The method according to any one of claims 1 to 3, characterized in that The splitting of the computation graph based on the type to obtain a plurality of first subgraphs includes: Determine, based on a mapping relationship, a fusion level between statements corresponding to the type, wherein the mapping relationship is used to describe a corresponding relationship between the type and the fusion level, and the fusion level is used to describe the fusibility of the statements; The computation graph is split based on the fusion level to obtain the multiple first sub-graphs.
5. The method according to claim 4, characterized in that The fusion level includes at least one of the following: the first level, the second level, the third level and the fourth level; the first level is used to indicate that the sentences can be fused front and back, the second level is used to indicate that the sentences cannot be fused back and forth, the third level is used to indicate that the sentences cannot be fused front and back, and the fourth level is used to indicate that the sentences may be fused front and back.
6. The method according to claim 4 or 5, characterized in that: The mapping relationship includes: the correspondence between the first level and the first type, the correspondence between the first level and the fifth type, the correspondence between the second level and the fourth type, the correspondence between the third level and the first combination type, the correspondence between the fourth level and the second type, the correspondence between the fourth level and the third type, and the correspondence between the fourth level and the sixth type; the first combination type includes the combination of the second type and the fourth type.
7. A data processing device, characterized in that: The data processing device comprises: An acquisition unit, used to acquire a computational graph, where the computational graph is used to represent the computational logic of multiple operator nodes; A modeling unit, configured to perform polyhedral modeling on the computation graph to obtain a plurality of statements; A splitting unit is used to split the computation graph based on the type of dependency relationship between the multiple statements to obtain multiple first subgraphs, wherein the dependency relationship is used to represent the sequential relationship of reading and / or writing between the multiple statements.
8. The data processing device according to claim 7, characterized in that The data processing device further comprises: An optimization unit is used to optimize the multiple first subgraphs to obtain multiple second subgraphs, the number of the multiple second subgraphs is greater than or equal to the number of the multiple first subgraphs, and the multiple second subgraphs are used to improve the processing performance of the data processing device.
9. The data processing device according to claim 7 or 8, characterized in that: The types include at least one of the following: a first type, a second type, a third type, a fourth type, a fifth type and a sixth type; the first type is used to indicate that the access order between statements is the same, and the number of instances of write operations and read operations between statements is the same; the second type is used to indicate that the access order between statements is different, and the number of instances of write operations and read operations between statements is the same; the third type is used to indicate that the number of write operations of each statement is greater than the number of read operations; the fourth type is used to indicate that the number of write operations of each statement is less than the number of read operations; the fifth type is used to indicate that the mapping relationship between statements is one-to-one; the sixth type is used to indicate that the mapping relationship between statements is not one-to-one.
10. The data processing device according to any one of claims 7 to 9, characterized in that: The splitting unit is specifically used to determine the fusion level between the statements corresponding to the type based on the mapping relationship, the mapping relationship is used to describe the corresponding relationship between the type and the fusion level, and the fusion level is used to describe the fusibility of the statements; The splitting unit is specifically used to split the computation graph based on the fusion level to obtain the multiple first subgraphs.
11. The data processing device according to claim 10, characterized in that The fusion level includes at least one of the following: the first level, the second level, the third level and the fourth level; the first level is used to indicate that the sentences can be fused front and back, the second level is used to indicate that the sentences cannot be fused back and forth, the third level is used to indicate that the sentences cannot be fused front and back, and the fourth level is used to indicate that the sentences may be fused front and back.
12. The data processing device according to claim 10 or 11, characterized in that The mapping relationship includes: the correspondence between the first level and the first type, the correspondence between the first level and the fifth type, the correspondence between the second level and the fourth type, the correspondence between the third level and the first combination type, the correspondence between the fourth level and the second type, the correspondence between the fourth level and the third type, and the correspondence between the fourth level and the sixth type; the first combination type includes the combination of the second type and the fourth type.
13. A data processing device, characterized in that: include: A processor, the processor is coupled to a memory, the memory is used to store programs or instructions, when the program or instructions are executed by the processor, the data processing device executes the method according to any one of claims 1 to 6.
14. A computer storage medium, characterized in that: The method comprises computer instructions, and when the computer instructions are executed on a terminal device, the terminal device executes the method according to any one of claims 1 to 6.
15. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.