Compiling method and device of operator, equipment and medium
By converting the Triton DSL operator source code into the Triton dialect and performing hardware-independent optimization, and then converting it into a hardware-specific dialect, and using the LLVM compilation tool to generate target code, the cross-platform problem of operator compilation under different chip architectures is solved, and cross-architecture operator versatility and performance improvement are achieved.
Patent Information
- Application Number
- CN202510725305.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-05-30
AI Technical Summary
When chips with different chip architectures are connected to Triton, there is a problem that the operators developed are difficult to use on chips with different chip architectures. This is mainly because chip manufacturers use closed-source compilers to protect architecture information, resulting in large version differences and difficulty in cross-architecture operator compilation.
By converting the operator source code written in the Triton domain-specific language into the Triton dialect, performing hardware-independent optimization, and then converting it into the hardware-specific dialect, a general programming language source code is generated and compiled using the LLVM compilation tool of the target chip, avoiding dependence on closed-source compilers.
It realizes universal compilation of operators on different chip architectures, improves compilation performance, protects chip architecture information from leakage, and supports the construction of cross-architecture operator libraries.
Smart Images

Figure CN120653254A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an operator compilation method, apparatus, device, and medium. Background Art
[0002] With the rise of generative artificial intelligence (AI), the competition for computing power is becoming increasingly fierce. Triton, a popular deep learning operator development tool based on Python, hides chip architecture information and leaves architecture-related optimizations to the compiler, significantly simplifying operator development.
[0003] However, when chips from manufacturers using different chip architectures are connected to Triton, in order to protect the architectural information of their chips, manufacturers usually connect to Triton in a closed-source code manner. Operators developed based on Triton need to be compiled entirely using the Triton compiler developed by the chip manufacturer. Due to the version differences of the Triton compilers from different manufacturers, it is difficult to develop operators that are universally applicable to chips with different chip architectures. Summary of the Invention
[0004] The present application provides an operator compilation method, apparatus, device, and medium, which enable operators to be universally applicable to chips with different chip architectures.
[0005] To achieve the above objectives, this application adopts the following technical solutions:
[0006] In a first aspect, a method for compiling an operator is provided, which may include: obtaining operator source code to be compiled, the operator source code being code written in a domain-specific language based on Triton; converting the operator source code into a Triton dialect for hardware-independent optimization, the Triton dialect being a hardware-independent dialect based on an MLIR framework, used to represent the computational logic of the operator source code; converting the Triton dialect into a hardware-specific dialect, the hardware-specific dialect being a hardware-related dialect based on an MLIR framework, used to represent hardware-oriented computational logic; generating source code in a general programming language based on the hardware-specific dialect; and calling an LLVM-based compilation tool corresponding to a target chip to generate target code based on the source code in a general programming language.
[0007] In one possible implementation, converting the Triton dialect into a hardware-specific dialect includes: converting the Triton dialect into a linear algebra dialect, which is a dialect used to represent linear algebra calculation operations in the MLIR framework; and converting the linear algebra dialect into the hardware-specific dialect.
[0008] In another possible implementation, converting the Triton dialect into a hardware-specific dialect further includes: converting arithmetic operations in the Triton dialect into an arithmetic dialect and / or a mathematical dialect, where the arithmetic dialect is a dialect used to represent mathematical operations in the MLIR framework, and the mathematical dialect is a dialect used to represent mathematical function operations in the MLIR framework; and converting the arithmetic dialect and / or the mathematical dialect into a hardware-specific dialect.
[0009] In another possible implementation, converting the Triton dialect into a hardware-specific dialect further includes: converting a tensor representation in the Triton dialect into a tensor dialect, where the tensor dialect is a dialect used to represent tensor operations in the MLIR framework; and converting the tensor dialect into a hardware-specific dialect.
[0010] In another possible implementation, converting the Triton dialect into a hardware-specific dialect further includes: converting the SCF in the Triton dialect into the SCF dialect, where the SCF dialect is a dialect used to represent SCF operations in the MLIR framework; and converting the SCF dialect into a hardware-specific dialect.
[0011] In another possible implementation, converting a linear algebra dialect into a hardware-specific dialect includes: converting the linear algebra dialect into an affine dialect and / or a vector dialect, where the affine dialect is a dialect used to represent affine transformations and loop operations in the MLIR framework, and the vector dialect is a dialect used to represent vector operations in the MLIR framework; and converting the affine dialect and / or the vector dialect into a hardware-specific dialect.
[0012] In another possible implementation, source code in a general programming language is generated according to the hardware-specific dialect, including: generating source code in C / C++ language based on the hardware-specific dialect using the EmitC code generation tool in the MLIR framework.
[0013] In a second aspect, a compilation device for an operator is provided, which may include: an acquisition module for acquiring operator source code to be compiled, the operator source code being code written in a domain-specific language based on Triton; a conversion module for converting the operator source code into a Triton dialect for hardware-independent optimization, the Triton dialect being a hardware-independent dialect based on the MLIR framework, used to represent the computational logic of the operator source code; converting the Triton dialect into a hardware-specific dialect, the hardware-specific dialect being a hardware-related dialect based on the MLIR framework, used to represent hardware-oriented computational logic; a generation module for generating source code in a general programming language based on the hardware-specific dialect; and calling the LLVM-based compilation tool corresponding to the target chip to generate target code based on the source code in the general programming language.
[0014] In one possible implementation, the conversion module is specifically configured to convert the Triton dialect into a linear algebra dialect, which is a dialect used to represent linear algebraic computational operations in the MLIR framework; and convert the linear algebra dialect into a hardware-specific dialect.
[0015] In another possible implementation, the conversion module is further used to convert arithmetic operations in the Triton dialect into an arithmetic dialect and / or a mathematical dialect, where the arithmetic dialect is a dialect used to represent mathematical operations in the MLIR framework, and the mathematical dialect is a dialect used to represent mathematical function operations in the MLIR framework; and convert the arithmetic dialect and / or the mathematical dialect into a hardware-specific dialect.
[0016] In another possible implementation, the conversion module is further used to convert the tensor representation in the Triton dialect into the tensor dialect, which is a dialect used to represent tensor operations in the MLIR framework; and convert the tensor dialect into a hardware-specific dialect.
[0017] In another possible implementation, the conversion module is further used to convert the SCF in the Triton dialect into the SCF dialect, where the SCF dialect is a dialect used to represent SCF operations in the MLIR framework; and convert the SCF dialect into a hardware-specific dialect.
[0018] In another possible implementation, the conversion module is specifically used to convert the linear algebra dialect into an affine dialect and / or a vector dialect, where the affine dialect is a dialect used to represent affine transformations and loop operations in the MLIR framework, and the vector dialect is a dialect used to represent vector operations in the MLIR framework; and convert the affine dialect and / or vector dialect into a hardware-specific dialect.
[0019] In another possible implementation, the generation module is specifically configured to generate source code in C / C++ language based on the hardware-specific dialect and the EmitC code generation tool in the MLIR framework.
[0020] In a third aspect, an electronic device is provided, comprising: a processor, a memory, and a communication interface. The memory and the communication interface are coupled to the processor, and the memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device performs the method described in any one of the first aspects.
[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, which, when executed on an electronic device, cause the electronic device to execute any one of the methods described in the first aspect.
[0022] In a fifth aspect, a computer program product comprising computer program instructions is provided, which, when executed by a processor, causes the processor to execute the method as described in any one of the above-mentioned first aspects.
[0023] In a sixth aspect, a device (for example, a chip system) is provided, comprising a processor for supporting an electronic device in implementing the method described in the first aspect above. In one possible design, the device further comprises a memory for storing program instructions and data necessary for the electronic device. When the device is a chip system, it may be composed of a chip or may include a chip and other discrete components.
[0024] In an embodiment of the present application, the operator source code to be compiled based on Triton can be first converted into the Triton dialect to perform hardware-independent optimization, thereby improving the performance of the target code finally compiled. Then, the Triton dialect can be converted into a hardware-oriented hardware-specific dialect related to hardware, so that the hardware-specific dialect can be converted into a source code using a general programming language, and then the target chip (i.e., the chip for running the operator) is called based on the compilation tool of the low-level virtual machine (Low Level Virtual Machine, LLVM), and the source code using the general programming language is compiled to generate the target code, thereby realizing hardware-related optimization of the operator and the generation of the target code. In this way, the compilation of the operator source code no longer needs to rely on the closed-source Triton compiler developed by the chip manufacturer, so that the developed operator can be universally connected to the chip of different chip architectures. Moreover, through the above-mentioned compilation method, since the final hardware-related compilation process is realized by the LLVM compilation tool of the manufacturer corresponding to the chip, the architecture information of the chip can still be protected by the chip manufacturer to avoid architecture information leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flowchart of a method for compiling an operator provided in an embodiment of the present application;
[0026] Figure 2 This is a flowchart of dialect conversion in a compilation method of an operator provided in an embodiment of the present application;
[0027] Figure 3 This is a schematic diagram of a code conversion process in a compilation method of an operator provided in an embodiment of the present application;
[0028] Figure 4 This is a schematic diagram of the structure of a compilation device for an operator provided in an embodiment of the present application;
[0029] Figure 5This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and appended claims of the present application, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that " / " means or, for example, A / B can mean A or B; "and / or" in the text is merely a description of an association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0031] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0032] The terms "first" and "second" in the following embodiments of this application are used for descriptive purposes only and should not be understood as implying or suggesting relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0033] With the rise of generative artificial intelligence (AI), the competition for computing power is becoming increasingly fierce. Triton, a popular deep learning operator development tool based on Python, hides chip architecture information and leaves architecture-related optimizations to the compiler, significantly simplifying operator development.
[0034] However, when chips from manufacturers using different chip architectures are connected to Triton, in order to protect the architectural information of their chips, manufacturers usually connect to Triton in a closed-source code manner. Operators developed based on Triton need to be compiled entirely using the Triton compiler developed by the chip manufacturer. Due to the version differences of the Triton compilers from different manufacturers, it is difficult to develop operators that are universally applicable to chips with different chip architectures.
[0035] In this regard, the present application provides a compilation method for an operator, for an operator source code (i.e., Triton DSL source code) written in a domain-specific language (DSL) based on Triton, using this method, the Triton DSL source code to be compiled can first be converted to obtain a Triton dialect (Dialect), so as to perform hardware-independent optimization based on Triton Dialect, thereby improving the performance of the target code finally compiled. Then, the Triton Dialect can be converted into a hardware-oriented hardware-related hardware-specific dialect (or called XPU-CDialect), so that the hardware-specific dialect can be converted into a source code using a general programming language, and then the target chip (i.e., the chip used to run the operator) is called based on a low-level virtual machine (LLVM) compilation tool to compile the above-mentioned source code using a general programming language to generate the target code, thereby achieving hardware-related optimization of the operator and generation of the target code. In this way, the compilation of the operator source code no longer needs to rely on the closed-source Triton compiler developed by the chip manufacturer, so that the developed operator can be universally connected to chips with different chip architectures. Moreover, through the above compilation method, since the final hardware-related compilation process is implemented through the chip manufacturer's own LLVM compilation tool, the chip architecture information can still be protected by the chip manufacturer to avoid architecture information leakage.
[0036] The following will describe in detail the operator compilation method provided by this application with reference to the accompanying drawings.
[0037] Reference Figure 1 , is a flow chart of a compilation method of an operator provided in an embodiment of the present application. Figure 1 As shown, the operator compilation method may specifically include S101-S105.
[0038] S101. Obtain the operator source code to be compiled.
[0039] The operator source code to be compiled may be a code written in a domain-specific language based on Triton, that is, an operator source code written in Triton DSL (or referred to as Triton DSL code).
[0040] Triton is a language and compiler for writing efficient parallel computing kernels, designed to provide optimized support for deep learning and high-performance computing tasks. It allows users to write high-performance operators for chips such as graphics processing units (GPUs) and neural network processors (NPUs) through a Python-based DSL, while using a multi-level intermediate representation (MLIR) framework to implement compilation. Therefore, in the embodiment of the present application, the operator source code to be compiled based on Triton DSL can have higher performance and can run better on the chip.
[0041] S102: Convert the operator source code into Triton Dialect and perform hardware-independent optimization.
[0042] The Triton Dialect is a hardware-independent dialect based on the MLIR framework that represents the computational logic of the corresponding operator source code. Specifically, the Triton Dialect represents the computational logic of the Triton DSL code by defining a set of specific operations (Ops), types (Types), and attributes (Attributes).
[0043] As a possible implementation method, when converting the operator source code to Triton Dialect, the operator source code can be first parsed into Python's Abstract Syntax Tree (AST). AST is a tree-like representation of the grammatical structure of the source code, which can represent the grammatical elements in the source code (such as variable declarations, function calls, expressions, etc.) as nodes in the tree structure. Therefore, it is possible to further capture the key information in the operator source code by traversing the AST, such as function definitions, variable declarations, operators, and memory access operations in the operator source code. Then, based on this key information, the operator source code can be converted into a Triton intermediate representation (IR). Triton IR is a hardware-independent high-level representation that can represent the computational logic of the operator source code. Thus, Triton IR can be converted into Triton Dialect based on the MLIR framework.
[0044] After converting the operator source code to Triton Dialect, hardware-independent optimizations (such as computational logic optimization) can be performed based on Triton Dialect to improve operator performance. For example, an inline expansion pass can be performed to expand sub-functions called in the kernel directly to the call point, reducing function call overhead and improving code efficiency. Another example is a combiner pass, which can be used to fuse operations on specific patterns to reduce redundant operations and improve code simplicity and efficiency. Another example is a canonicalizer pass, which simplifies and standardizes the code using a series of rules to eliminate redundant expressions, making the Triton Dialect more standardized and facilitating subsequent optimization. Another example is a common subexpression elimination pass (CSE pass), which identifies and eliminates common subexpressions to avoid repeated calculations and improve code execution efficiency. Another example is a loop-invariant code motion pass (LICM pass), which moves loop-independent code out of the loop body, reducing the amount of computation within the loop and improving loop efficiency. Of course, the above is only an example of hardware-independent optimization based on Triton Dialect. In some other possible implementations of this application, some other hardware-independent optimizations can also be performed, which are not limited here.
[0045] S103, converting the Triton Dialect into a hardware-specific dialect.
[0046] The hardware-specific dialect (also known as the XPU-C dialect) is a hardware-specific dialect based on the MLIR framework, used to represent hardware-oriented computing logic. By defining specific operations, types, and attributes, the hardware-specific dialect provides a representation of operations that are closely related to the hardware, directly supporting hardware features and facilitating subsequent hardware-related optimization of the code.
[0047] Figure 2 A schematic diagram of the process of dialect conversion in a compilation method of an operator provided in an embodiment of the present application is shown.
[0048] As a possible implementation method, refer to Figure 2 As shown, when converting Triton Dialect to a hardware-specific dialect, Triton Dialect can be first converted to Linalg Dialect (linear algebra dialect), and then Linalg Dialect can be converted to the corresponding hardware-specific dialect.
[0049] Linalg Dialect is a dialect used in the MLIR framework to represent linear algebra operations. This dialect enables advanced linear algebra operations on tensors and supports various tensor shapes, data types, memory layouts, and conversions. Linalg Dialect is lower-level than Triton Dialect, so converting from Triton Dialect to Linalg Dialect and then to a hardware-specific dialect is more convenient and reliable.
[0050] For example, based on the triton-shared-opt tool (Triton shared optimization tool), the Triton Dialect can be converted into the corresponding Linalg Dialect through a triton-to-linalg conversion pass (ie, Triton to Linalg conversion optimization).
[0051] As an example, when converting Linalg Dialect to a hardware-specific dialect, such as Figure 2 As shown, the Linalg Dialect can also be converted into a dialect that is closer to the hardware, such as Affine Dialect and / or Vector Dialect, and then these dialects that are closer to the hardware can be further converted into hardware-specific dialects. Among them, Affine Dialect is the dialect used to represent affine transformations and loop operations in the MLIR framework, and Vector Dialect is the dialect used to represent vector operations in the MLIR framework. Therefore, the affine transformations and loop operations in the Linalg Dialect can be converted into Affine Dialect, and the vector operations in the Linalg Dialect can be converted into Vector Dialect. Since these dialects are closer to the hardware, first converting the Linalg Dialect into a dialect that is closer to the hardware and then further converting it into a hardware-specific dialect can improve the conversion effect and thus improve the code performance. Of course, in other possible implementations of the present application, the Linalg Dialect can also be directly converted into a hardware-specific dialect, or converted into other intermediate dialects and then further converted into a hardware-specific dialect, which is not limited here.
[0052] As another possible implementation, if the Triton Dialect contains arithmetic operations, refer to Figure 2As shown, when converting Triton Dialect to a hardware-specific dialect, the arithmetic operations in Triton Dialect can also be converted to Arith Dialect (arithmetic dialect) and / or Math Dialect (mathematical dialect). Arith Dialect and / or Math Dialect are then converted to hardware-specific dialects. Among them, Arith Dialect is a dialect used to represent mathematical arithmetic operations in the MLIR framework, which can define basic integer and floating-point mathematical operations. Math Dialect is a dialect used to represent mathematical function operations in the MLIR framework, which can support the implementation of more complex mathematical function operations. Therefore, the basic mathematical arithmetic operations in Triton Dialect can be converted to Arith Dialect based on the optimization tool chain of the MLIR framework, and the complex mathematical function operations in Triton Dialect can be converted to Math Dialect. In this way, the arithmetic operations in Triton Dialect can be converted to a lower-level dialect, so that the arithmetic operations in Triton Dialect can be further converted to hardware-specific dialects more conveniently and efficiently.
[0053] As another possible implementation, if the Triton Dialect includes tensor representation, refer to Figure 2 As shown, when converting Triton Dialect to a hardware-specific dialect, the tensor representation in Triton Dialect can also be converted to Tensor Dialect (tensor dialect). Then convert Tensor Dialect to a hardware-specific dialect. Among them, TensorDialect is a dialect used to represent tensor operations in the MLIR framework, which can define a series of operations for creating, operating, and converting tensors. Therefore, the tensor representation in Triton Dialect can be converted to TensorDialect based on the optimization tool chain of the MLIR framework, thereby converting the tensor representation in Triton Dialect to a lower-level dialect, so that the tensor representation in Triton Dialect can be further converted to a hardware-specific dialect more conveniently and efficiently.
[0054] As another possible implementation, if the Triton Dialect includes Structured Control Flow (SCF), refer to Figure 2As shown, when converting Triton Dialect into a hardware-specific dialect, the SCF in Triton Dialect can also be converted into SCF Dialect (Structured Control Flow Dialect). Then convert SCFDialect into a hardware-specific dialect. Among them, SCF Dialect is a dialect used to represent structured control flow operations in the MLIR framework, which can express operations of common control flow structures, such as loops, conditional statements, and parallel execution. Therefore, the structured control flow in Triton Dialect can be converted into SCF Dialect based on the optimization tool chain of the MLIR framework, thereby converting the structured control flow in Triton Dialect into a lower-level dialect, so that the structured control flow in Triton Dialect can be further converted into a hardware-specific dialect more conveniently and efficiently.
[0055] In the examples of this application, refer to Figure 2 As shown, you can also extend the representation of specific computational logic by using custom dialects (such as XCSExt Dialect), thereby converting these specific computational logics in Triton Dialect into corresponding custom dialects, so as to better convert certain specific computational logics in Triton Dialect into hardware-specific dialects. In actual applications, you can set it according to the specific situation and there is no restriction here.
[0056] For example, in practical applications, such as Figure 2 As shown, the vector operations in the Affine Dialect can be further converted to the Vector Dialect that is closer to the hardware, and the structured control flow in the Affine Dialect can be further converted to the SCF Dialect that is closer to the hardware in order to better convert the Affine Dialect to the hardware-specific dialect.
[0057] S104: Generate source code in a general programming language according to the hardware-specific dialect.
[0058] The general programming language used in the generated source code can be the same as the general programming language corresponding to the LLVM compilation tool of the back-end of the target chip, so as to facilitate the subsequent compilation of the source code according to the LLVM compilation tool corresponding to the target chip.
[0059] As an example, the EmitC (or MLIR EmitC) code generation tool under the MLIR framework provides a set of operations and types that can convert the dialect in the MLIR framework into portable C / C++ language source code. Therefore, in this example, the hardware-specific dialect can be converted into C / C++ source code by the MLIR EmitC code generation tool. Of course, in other possible implementations of the present application, other code generation tools can also be used to generate corresponding source code in a general programming language based on the hardware-specific dialect, which is not limited here.
[0060] S105 , calling the LLVM-based compilation tool corresponding to the target chip, and generating target code according to the source code in the general programming language.
[0061] The corresponding LLVM compilation tool can be called through the interface provided by the manufacturer of the target chip. The target code can be machine code that can be executed by the target chip.
[0062] Since chip manufacturers usually hide the chip's architecture information in the form of private LLVM IR in the LLVM-based compilation tools they provide in order to protect the chip's architecture information. Therefore, in this application, by calling the LLVM-based compilation tool corresponding to the target chip (i.e., the LLVM compilation tool provided by the target chip manufacturer), the LLVM compilation tool of the target chip is used to compile the previously generated source code in a general programming language (such as C / C++ source code). This allows hardware-related optimization of operators without leaking the chip's architecture information, and ultimately generates the target code.
[0063] Based on the aforementioned implementation, the operator compilation method provided by this application can gradually convert the operator source code based on Triton DSL into target code in stages, and can perform hardware-independent optimization on the operator source code to improve the performance of the operator. For example, Figure 3 A schematic diagram of the code conversion process in a compilation method of an operator provided in an embodiment of the present application.
[0064] like Figure 3 As shown, the compilation process of the operator can be divided into five stages, and the codes corresponding to each stage are as follows Figure 3As shown. The code corresponding to the first stage is Triton Dialect code. This stage converts the operator source code based on Triton DSL into Triton Dialect for hardware-independent optimization. The code corresponding to the second stage is Linalg Dialect code. This stage converts the Triton Dialect code into Linalg Dialect code so that it can be more conveniently and accurately converted to hardware-specific dialect code. The code corresponding to the third stage is hardware-specific code. This stage converts the Linalg Dialect code into a hardware-specific dialect for subsequent hardware-related optimization. The code corresponding to the fourth stage is C / C++ source code. This stage converts the hardware-specific dialect into portable C / C++ source code using the MLIR EmitC code generation tool under the MLIR framework. This allows the subsequent fifth stage to call the LLVM compilation tool corresponding to the target chip based on this source code to perform hardware-related optimization on the operator and generate target code. Accordingly, the code corresponding to the fifth stage is the final target code.
[0065] The operator compilation method provided in the embodiment of the present application can convert the operator source code based on Triton DSL into Triton Dialect to optimize the operator in a hardware-independent manner and improve the operator performance. Then, by converting Triton Dialect into a hardware-specific dialect for hardware, and then converting the hardware-specific dialect into the source code of a general programming language, the LLVM compilation tool corresponding to the target chip is called to compile and optimize the source code to obtain the target code, thereby protecting the architecture information of the target chip and avoiding architecture information leakage without relying on the dedicated Triton compiler provided by the manufacturer of the target chip. Based on this, the operator compilation method provided in the embodiment of the present application can realize the compilation of operators that are common to chips of different architectures, so that operators written based on Triton DSL can be compiled and used in chips of different architectures, thereby facilitating the construction of an operator library that is common to chips of different architectures.
[0066] It is obvious that those skilled in the art can make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0067] It is understandable that, in order to implement the above functions, the electronic device may include hardware structures and / or software modules that perform the corresponding functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0068] The embodiment of the present application can divide the functional modules according to the above method examples, thereby obtaining a device that can implement the above method examples. For example, Figure 4 FIG. 1 shows a schematic diagram of the structure of a compilation device for an operator provided in an embodiment of the present application. Figure 4 As shown, the device may include: an acquisition module 401, used to obtain an operator source code to be compiled, where the operator source code is a code written in a domain-specific language based on Triton; a conversion module 402, used to convert the operator source code into a Triton dialect for hardware-independent optimization, where the Triton dialect is a hardware-independent dialect based on the MLIR framework, used to represent the computational logic of the operator source code; converting the Triton dialect into a hardware-specific dialect, where the hardware-specific dialect is a hardware-related dialect based on the MLIR framework, used to represent hardware-oriented computational logic; a generation module 403, used to generate a source code in a general programming language based on the hardware-specific dialect; and calling an LLVM-based compilation tool corresponding to the target chip to generate a target code based on the source code in a general programming language.
[0069] In one possible implementation, the conversion module 402 is specifically configured to convert the Triton dialect into a linear algebra dialect, which is a dialect used to represent linear algebraic computational operations in the MLIR framework; and convert the linear algebra dialect into a hardware-specific dialect.
[0070] In another possible implementation, the conversion module 402 is further used to convert arithmetic operations in the Triton dialect into an arithmetic dialect and / or a mathematical dialect, where the arithmetic dialect is a dialect used to represent mathematical operations in the MLIR framework, and the mathematical dialect is a dialect used to represent mathematical function operations in the MLIR framework; and convert the arithmetic dialect and / or the mathematical dialect into a hardware-specific dialect.
[0071] In another possible implementation, the conversion module 402 is further configured to convert the tensor representation in the Triton dialect into a tensor dialect, where the tensor dialect is a dialect used to represent tensor operations in the MLIR framework; and convert the tensor dialect into a hardware-specific dialect.
[0072] In another possible implementation, the conversion module 402 is further configured to convert the SCF in the Triton dialect into the SCF dialect, where the SCF dialect is a dialect used to represent SCF operations in the MLIR framework; and convert the SCF dialect into a hardware-specific dialect.
[0073] In another possible implementation, the conversion module 402 is specifically configured to convert the linear algebra dialect into an affine dialect and / or a vector dialect, where the affine dialect is a dialect used to represent affine transformations and loop operations in the MLIR framework, and the vector dialect is a dialect used to represent vector operations in the MLIR framework; and convert the affine dialect and / or the vector dialect into a hardware-specific dialect.
[0074] In another possible implementation, the generation module 403 is specifically configured to generate source code in C / C++ language based on the hardware-specific dialect and the EmitC code generation tool in the MLIR framework.
[0075] The embodiment of the present application also provides an electronic device, Figure 5 FIG. 1 shows a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device includes: a bus 501, a processor 502, a memory 503, and a communication interface 504. The processor 502, the memory 503, and the communication interface 504 communicate with each other via the bus 501. The memory 503 stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor 502, the electronic device executes the method provided in the above embodiment. The electronic device can be a server or a terminal device. It should be understood that this application does not limit the number of processors 502 and memories 503 in the electronic device.
[0076] The bus 501 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus 501 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5The bus 501 may include a path for transmitting information between various components of the electronic device (eg, memory 503, processor 502, communication interface 504).
[0077] The processor 502 may include any one or more processors such as a central processing unit, a graphics processing unit, a microprocessor (MP), or a digital signal processor (DSP).
[0078] The memory 503 may include a volatile memory, such as a random access memory (RAM). The memory 503 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0079] The communication interface 504 uses a command distribution module such as, but not limited to, a network interface card, a transceiver, etc. to implement communication between the electronic device and other devices or a communication network.
[0080] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the method provided in the aforementioned embodiment.
[0081] In addition, in the embodiments of the present application, the functional units or modules of the apparatus for implementing the aforementioned method examples may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0082] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0083] Therefore, an embodiment of the present application further provides a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, the processor executes the method provided in the aforementioned embodiment.
[0084] An embodiment of the present application further provides a chip system comprising at least one processor and at least one interface circuit. The processor and the interface circuit may be interconnected via a circuit. For example, the interface circuit may be used to receive signals from another device (e.g., a memory of an electronic device). For another example, the interface circuit may be used to send signals to another device (e.g., a processor).
[0085] For example, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the chip system can perform the various steps in the above embodiment. Of course, the chip system can also include other discrete components, which are not specifically limited in the embodiments of the present application.
[0086] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for compiling an operator, characterized in that: include: Obtain the operator source code to be compiled, where the operator source code is written in a domain-specific language based on Triton; Convert the operator source code into the Triton dialect for hardware-independent optimization. The Triton dialect is a hardware-independent dialect based on the Multi-Level Intermediate Representation (MLIR) framework, used to represent the computational logic of the operator source code. Converting the Triton dialect into a hardware-specific dialect, where the hardware-specific dialect is a hardware-related dialect based on the MLIR framework and is used to represent hardware-oriented computing logic; generating source code in a general programming language according to the hardware-specific dialect; A compilation tool based on the low-level virtual machine LLVM corresponding to the target chip is called to generate target code according to the source code using the general programming language.
2. The method according to claim 1, characterized in that The conversion of the Triton dialect to a hardware-specific dialect includes: Converting the Triton dialect into a linear algebra dialect, where the linear algebra dialect is a dialect used to represent linear algebra operations in the MLIR framework; Convert the linear algebra dialect to the hardware-specific dialect.
3. The method according to claim 2, characterized in that Converting the Triton dialect to a hardware-specific dialect further includes: Converting arithmetic operations in the Triton dialect into an arithmetic dialect and / or a mathematical dialect, wherein the arithmetic dialect is a dialect used to represent mathematical arithmetic operations in the MLIR framework, and the mathematical dialect is a dialect used to represent mathematical function operations in the MLIR framework; The arithmetic dialect and / or the mathematical dialect is converted into the hardware specific dialect.
4. The method according to claim 2, characterized in that Converting the Triton dialect to a hardware-specific dialect further includes: Converting the tensor representation in the Triton dialect to a tensor dialect, where the tensor dialect is a dialect used to represent tensor operations in the MLIR framework; Convert the tensor dialect to the hardware-specific dialect.
5. The method according to claim 2, characterized in that Converting the Triton dialect to a hardware-specific dialect further includes: Converting the structured control flow (SCF) in the Triton dialect to the SCF dialect, where the SCF dialect is a dialect for representing SCF operations in the MLIR framework. The SCF dialect is converted to the hardware specific dialect.
6. The method according to claim 2, characterized in that The converting the linear algebra dialect into the hardware specific dialect comprises: Convert the linear algebra dialect into an affine dialect and / or a vector dialect, wherein the affine dialect is a dialect used to represent affine transformations and loop operations in the MLIR framework, and the vector dialect is a dialect used to represent vector operations in the MLIR framework; The affine dialect and / or the vector dialect is converted to the hardware specific dialect.
7. The method according to any one of claims 1 to 6, characterized in that Generating source code in a general programming language according to the hardware-specific dialect includes: According to the hardware-specific dialect, source code in the C / C++ language is generated based on the EmitC code generation tool in the MLIR framework.
8. A compilation device for an operator, characterized in that: include: An acquisition module is used to obtain the operator source code to be compiled. The operator source code is a code written in a domain-specific language based on Triton; a conversion module for converting the operator source code into a Triton dialect for hardware-independent optimization, wherein the Triton dialect is a hardware-independent dialect based on the multi-level intermediate representation (MLIR) framework and is used to represent the computational logic of the operator source code; and converting the Triton dialect into a hardware-specific dialect, wherein the hardware-specific dialect is a hardware-dependent dialect based on the MLIR framework and is used to represent hardware-oriented computational logic; a generation module, configured to generate source code in a general programming language according to the hardware-specific dialect; A compilation tool based on the low-level virtual machine LLVM corresponding to the target chip is called to generate target code according to the source code using the general programming language.
9. An electronic device, characterized in that: The electronic device includes: a processor, a memory and a communication interface; the memory and the communication interface are coupled to the processor, the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for compiling and optimizing heterogeneous platform based on high-order cryptographic operator
CN116301894A
Multi-core heterogeneous computing framework and engineering machinery
CN116382705A
Deep learning model compiler based on MLIR
CN117332850A
Deep learning model compiling method and device, electronic equipment and storage medium
CN118409756A
Triton compiler assembly line-oriented optimization system and optimization method
CN118605850A