Compiling method and device of operator, equipment and medium

By converting the operator source code into Triton intermediate representation for universal optimization and compiling it using the Triton compiler of the target chip, the problem of operator compilation being unable to be open-sourced for optimization under the protection of chip architecture information is solved, and compilation performance and efficiency are improved without leaking chip information.

CN120653253APending Publication Date: 2025-09-16ZHONGKE JIAHE (HANGZHOU) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510718440.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Currently, when chips with different chip architectures are connected to Triton, they are usually connected to the Triton compiler in a closed-source manner for the purpose of protecting chip architecture information, resulting in the operator compilation process being unable to be optimized and expanded based on the open source community.

Method used

The operator source code is converted into Triton intermediate representation, and optimization that does not involve chip information is performed. After general optimization using Triton intermediate representation, it is compiled through the Triton compiler corresponding to the target chip. It is divided into two parts: general optimization and closed-source proprietary optimization, realizing partial open source compilation.

Benefits of technology

Without leaking chip information, the operator compilation performance is improved, allowing part of the compilation process to be optimized and expanded based on the open source community, thereby improving the compilation efficiency and performance of the operator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653253A_ABST
    Figure CN120653253A_ABST
Patent Text Reader

Abstract

The invention provides an operator compiling method and device, equipment and a medium, and relates to the technical field of artificial intelligence. The operator compiling method comprises the following steps: obtaining a to-be-compiled operator source code, wherein the operator source code is a code compiled by adopting a field-specific language based on Triton; converting the operator source code into Triton intermediate representation, wherein the Triton intermediate representation is used for representing computational logic of the operator source code; carrying out optimization which does not involve chip information on an operator based on the Triton intermediate representation to obtain an optimized intermediate representation; and compiling the optimized intermediate representation through a Triton compiler corresponding to the target chip to obtain a target code. Open sources can be conveniently realized in partial stages in the operator compiling process, so that the partial compiling process of the operator can be conveniently optimized and expanded based on the open sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an operator compilation method, apparatus, device, and medium. Background Art

[0002] With the rise of generative artificial intelligence (AI), the competition for computing power is becoming increasingly fierce. Triton, a popular deep learning operator development tool based on Python, hides chip architecture information and leaves architecture-related optimizations to the compiler, significantly simplifying operator development.

[0003] However, when chips from manufacturers using different chip architectures are connected to Triton, in order to protect the architectural information of their chips, manufacturers usually connect to Triton in a closed-source code manner. Operators developed based on Triton need to be compiled entirely using the Triton compiler developed by the chip manufacturer, resulting in the operator compilation process being unable to be optimized and expanded based on the open source community. Summary of the Invention

[0004] The present application provides an operator compilation method, apparatus, device, and medium, which can facilitate the open source implementation of some stages in the operator compilation process, thereby facilitating the optimization and expansion of part of the operator compilation process based on open source.

[0005] To achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, a method for compiling an operator is provided, which may include: obtaining operator source code to be compiled, the operator source code being code written in a domain-specific language based on Triton; converting the operator source code into a Triton intermediate representation, which is used to represent the computational logic of the operator source code; optimizing the operator based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation; and compiling and processing the optimized intermediate representation through the Triton compiler corresponding to the target chip to obtain a target code.

[0007] In one possible implementation, the target chip is a general-purpose graphics processor, and the operator is optimized based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation, including: converting the Triton intermediate representation into the Triton GPU Intermediate Representation (TTGIR), which is an intermediate representation based on the Multi-Level Intermediate Representation (MLIR) framework, used to express the computational logic of the operator source code on the graphics processor; optimizing the operator based on the TTGIR without involving chip information to obtain an optimized intermediate representation.

[0008] In another possible implementation, the target chip is an application-specific integrated circuit, and the operator is optimized based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation, including: converting the Triton intermediate representation into a linear algebra dialect, which is an intermediate representation used to represent linear algebraic calculation operations in the MLIR framework; and optimizing the operator based on the linear algebra dialect without involving chip information to obtain an optimized intermediate representation.

[0009] In another possible implementation, converting the Triton intermediate representation into a linear algebra dialect includes: converting the Triton intermediate representation into a Triton dialect, where the Triton dialect is a dialect used to represent computational logic in the MLIR framework; and converting the Triton dialect into a linear algebra dialect.

[0010] In another possible implementation, the operator source code is converted into a Triton intermediate representation, including: parsing the operator source code into a Python abstract syntax tree; traversing the abstract syntax tree to obtain information nodes in the operator source code; and converting the operator source code into a Triton intermediate representation based on the information nodes.

[0011] In another possible implementation, the chip information includes at least one of the chip's architecture type, the number of cores of the chip, the chip's instruction set extension architecture information, the chip's memory hierarchy information, the chip's memory type, the chip's bus architecture information, and the chip's interface information.

[0012] In another possible implementation, before compiling the optimized intermediate representation through the Triton compiler corresponding to the target chip to obtain the target code, the method further includes: calling the Triton compiler through the compilation interface corresponding to the target chip.

[0013] On the second aspect, a compilation device for an operator is provided, which may include: an acquisition module for acquiring the operator source code to be compiled, the operator source code being a code written in a domain-specific language based on Triton; a conversion module for converting the operator source code into a Triton intermediate representation, the Triton intermediate representation being used to represent the computational logic of the operator source code; an optimization module for optimizing the operator based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation; a compilation module for compiling and processing the optimized intermediate representation through the Triton compiler corresponding to the target chip to obtain a target code.

[0014] In one possible implementation, the target chip is a general-purpose graphics processor, and the optimization module is specifically used to convert the Triton intermediate representation into TTGIR, which is an intermediate representation based on the MLIR framework and is used to express the computational logic of the operator source code on the graphics processor; based on TTGIR, the operator is optimized without involving chip information to obtain an optimized intermediate representation.

[0015] In another possible implementation, the target chip is an application-specific integrated circuit, and the optimization module is specifically used to convert the Triton intermediate representation into a linear algebra dialect, which is an intermediate representation used to represent linear algebra calculation operations in the MLIR framework; based on the linear algebra dialect, the operator is optimized without involving chip information to obtain an optimized intermediate representation.

[0016] In another possible implementation, the optimization module is specifically used to convert the Triton intermediate representation into the Triton dialect, which is a dialect used to represent computational logic in the MLIR framework; and convert the Triton dialect into a linear algebra dialect.

[0017] In another possible implementation, the conversion module is specifically used to parse the operator source code into a Python abstract syntax tree; traverse the abstract syntax tree to obtain information nodes in the operator source code; and based on the information nodes, convert the operator source code into a Triton intermediate representation.

[0018] In another possible implementation, the chip information includes at least one of the chip's architecture type, the number of cores of the chip, the chip's instruction set extension architecture information, the chip's memory hierarchy information, the chip's memory type, the chip's bus architecture information, and the chip's interface information.

[0019] In another possible implementation, the compilation module is further configured to call the Triton compiler through a compilation interface corresponding to the target chip.

[0020] In a third aspect, an electronic device is provided, comprising: a processor, a memory, and a communication interface. The memory and the communication interface are coupled to the processor, and the memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device performs the method described in any one of the first aspects.

[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, which, when executed on an electronic device, cause the electronic device to execute any one of the methods described in the first aspect.

[0022] In a fifth aspect, a computer program product comprising computer program instructions is provided, which, when executed by a processor, causes the processor to execute the method as described in any one of the above-mentioned first aspects.

[0023] In a sixth aspect, a device (for example, a chip system) is provided, comprising a processor for supporting an electronic device in implementing the method described in the first aspect above. In one possible design, the device further comprises a memory for storing program instructions and data necessary for the electronic device. When the device is a chip system, it may be composed of a chip or may include a chip and other discrete components.

[0024] In an embodiment of the present application, the operator source code to be compiled based on Triton can be first converted into a Triton intermediate representation, so that optimization (or hardware-independent optimization, general optimization, etc.) not involving chip information can be performed based on the Triton intermediate representation, thereby improving the performance of the target code finally compiled. Then, the optimized intermediate representation can be compiled based on the Triton compiler corresponding to the target chip (i.e., the chip for running the operator) to generate target code, thereby introducing the chip information of the target chip to perform hardware-related optimization on the operator and generate target code. In this way, the compilation process of the operator source code can be divided into a general optimization (i.e., optimization not involving chip information) part based on the Triton intermediate representation, and a closed-source proprietary optimization part based on the introduction of chip information of the Triton compiler corresponding to the target chip. Since the general optimization part based on the Triton intermediate representation does not involve chip information, this part of the compilation process can be open source, thereby facilitating optimization and expansion of this part of the compilation process based on the open source community. And the closed-source proprietary optimization part of the Triton compiler corresponding to the target chip, its Triton compiler can be provided by the manufacturer of the target chip, and the manufacturer can still adopt a closed-source approach to avoid leakage of its chip information. This operator compilation method can open source some stages of the operator compilation process without leaking chip information, so as to optimize and expand the operator compilation process based on the open source community and improve the operator compilation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flowchart of a method for compiling an operator provided in an embodiment of the present application;

[0026] Figure 2 This is a flowchart of another operator compilation method provided in an embodiment of the present application;

[0027] Figure 3 This is a schematic diagram of the structure of a compilation device for an operator provided in an embodiment of the present application;

[0028] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and appended claims of the present application, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that " / " means or, for example, A / B can mean A or B; "and / or" in the text is merely a description of an association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0030] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0031] The terms "first" and "second" in the following embodiments of this application are used for descriptive purposes only and should not be understood as implying or suggesting relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0032] With the rise of generative artificial intelligence (AI), the competition for computing power is becoming increasingly fierce. Triton, a popular deep learning operator development tool based on Python, hides chip architecture information and leaves architecture-related optimizations to the compiler, significantly simplifying operator development.

[0033] However, when chips from manufacturers using different chip architectures are connected to Triton, in order to protect the architectural information of their chips, manufacturers usually connect to Triton in a closed-source code manner. Operators developed based on Triton need to be compiled entirely using the Triton compiler developed by the chip manufacturer, resulting in the operator compilation process being unable to be optimized and expanded based on the open source community.

[0034] In this regard, the present application provides a compilation method for an operator, for an operator source code (i.e., Triton DSL source code) written in a domain-specific language (DSL) based on Triton, using this method, the Triton DSL source code to be compiled can be first converted to obtain a Triton intermediate representation, so as to perform general optimization based on the Triton intermediate representation that does not involve information, thereby improving the performance of the target code finally compiled. Then, the optimized intermediate representation can be compiled based on the Triton compiler corresponding to the target chip (i.e., the chip for running the operator) to generate the target code, thereby introducing the chip information of the target chip to perform hardware-related optimization on the operator and generate the target code. In this way, the compilation process of the operator source code can be divided into a general optimization based on the Triton intermediate representation (i.e., optimization that does not involve chip information) part, and a closed-source proprietary optimization part based on the Triton compiler corresponding to the target chip that introduces chip information. Since the general optimization part based on the Triton intermediate representation does not involve chip information, this part of the compilation process can be open source, thereby facilitating optimization and expansion of this part of the compilation process based on the open source community. The Triton compiler for the target chip can be provided by the target chip manufacturer, which can still adopt a closed-source approach to prevent chip information leakage. This operator compilation method can open source some stages of the operator compilation process without leaking chip information, allowing the open source community to optimize and expand the operator compilation process and improve operator compilation performance.

[0035] The following will describe in detail the operator compilation method provided by this application with reference to the accompanying drawings.

[0036] Reference Figure 1 , is a flow chart of a compilation method of an operator provided in an embodiment of the present application. Figure 1 As shown, the compilation method of the operator may specifically include S101-S104.

[0037] S101. Obtain operator source code to be compiled.

[0038] The operator source code to be compiled may be a code written in a domain-specific language based on Triton, that is, an operator source code written in Triton DSL (or referred to as Triton DSL code).

[0039] Triton is a language and compiler for writing efficient parallel computing kernels, designed to provide optimized support for deep learning and high-performance computing tasks. It allows users to write high-performance operators for chips such as graphics processing units (GPUs) and neural network processors (NPUs) through a Python (a high-level programming language with concise syntax and powerful libraries)-based DSL, while using a multi-level intermediate representation (MLIR) framework to implement compilation. Therefore, in the embodiment of the present application, the operator source code to be compiled based on Triton DSL can have higher performance and can run better on the chip.

[0040] S102: Convert the operator source code into Triton intermediate representation.

[0041] The Triton Intermediate Representation (IR) is a hardware-independent intermediate representation. It can be used to represent the computational logic of an operator's source code, describing the logic and structure of the operator while abstracting away specific hardware details. This allows for more convenient operator optimization based on the Triton IR. Specifically, the Triton IR provides high-level abstractions, expressing tensor operations in the form of parameterized block variables.

[0042] As a possible implementation method, when converting the operator source code into the Triton intermediate representation, the operator source code can be first parsed into Python's abstract syntax tree (AST). AST is a tree representation of the grammatical structure of the source code, which can represent the grammatical elements (such as variable declarations, function calls, expressions, etc.) in the source code as nodes in the tree structure. Therefore, it is possible to further capture the key information (or information nodes) in the operator source code by traversing the AST, for example, function definitions, variable declarations, operators, and memory access operations in the operator source code. Then, based on these key information, the operator source code can be converted into the Triton intermediate representation. Of course, the above is only an exemplary description. In some other possible implementations, the operator source code can also be converted into the Triton intermediate representation by other means, which are not limited here.

[0043] S103. Optimize the operator without involving chip information based on the Triton intermediate representation to obtain an optimized intermediate representation.

[0044] Among them, chip information is the architecture information of the chip. For example, the chip information may include the chip architecture type (such as reduced instruction set computer architecture, complex instruction set computer architecture, etc.), the number of cores of the chip, the chip instruction set extension architecture (such as single instruction multiple data stream architecture, multiple instruction multiple data stream architecture, etc.), the chip memory hierarchy information, the chip memory type (such as static random access memory, dynamic random access memory, etc.), the chip bus architecture information, the chip interface information, etc.

[0045] For example, the Triton intermediate representation can be further converted and downgraded, so that it can be optimized without involving chip information during the conversion process, and finally a hardware-specific dialect for hardware (i.e., the optimized intermediate representation) can be obtained. Among them, the hardware-specific dialect is an intermediate representation based on the MLIR framework that is related to the hardware and can introduce hardware characteristics to represent hardware-oriented computing logic. The hardware-specific dialect can provide an operation representation that is closely related to the hardware by defining specific operations, types, and attributes to directly support hardware characteristics, thereby facilitating subsequent hardware-related optimization of the code.

[0046] Of course, different Triton intermediate representation optimization processes can also be used based on different types of target chips (i.e., chips used to run operators), so as to better optimize operators based on Triton intermediate representation without involving chip information.

[0047] As an example, when the target chip is a general-purpose graphics processing unit (GPGPU), the Triton intermediate representation can be first converted to the Triton GPU Intermediate Representation (TTGIR). Then, based on TTGIR, the operator can be optimized without chip information to obtain the optimized intermediate representation. TTGIR is an intermediate representation based on the MLIR framework. Compared to the Triton intermediate representation, it incorporates GPU hardware features such as thread configuration, memory allocation, and asynchronous operations. TTGIR also adds some operations specifically for GPUs, such as the operation for allocating shared memory (alloc_tensor) and the operation for asynchronously inserting data slices (insert_slice_async). TTGIR can provide a representation that is closer to the GPU hardware, allowing code optimization and better utilization of the GPU's parallel computing capabilities. For example, the Triton intermediate representation can be converted to TTGIR based on the add_convert_to_ttgpuir conversion pass (i.e., the conversion optimization process from Triton intermediate representation to TTGIR).

[0048] Among them, the performance of operators can be improved through optimization based on TTGIR that does not involve chip information. For example, Common Subexpression Elimination (CSE) can be performed to identify and eliminate repeated subexpressions and reduce unnecessary calculations. For another example, Dead Code Elimination (DCE) can be performed to remove unused variables, operations, or code fragments. For another example, Instruction Reordering can be performed to optimize the execution order of instructions and reduce dependency delays. For another example, Loop Fusion can be performed to merge multiple loops into one loop to reduce loop overhead. For another example, Constant Propagation can be performed to propagate constant values ​​into expressions to reduce runtime calculations. For another example, Memory Access Pattern Optimization can be performed to optimize memory access patterns and reduce cache misses. For another example, Layout Normalization can be performed to convert data layout into a standard form to facilitate subsequent optimization. Of course, the above is only an example of TTGIR-based optimization that does not involve chip information. In some other possible implementations of the present application, some other optimizations that do not involve chip information can also be performed, which is not limited here.

[0049] As another example, when the target chip is an application-specific integrated circuit (ASIC), the Triton intermediate representation can be first converted to a linear algebra dialect (LinalgDialect), and then the operator can be optimized based on the Linalg Dialect without chip information to obtain the optimized intermediate representation. Among them, Linalg Dialect is a dialect used to represent linear algebra operations in the MLIR framework. This dialect can be used to process high-level linear algebra operations on tensors and supports various tensor shapes, data types, memory layouts, and conversions. The level of Linalg Dialect is lower than that of the Triton intermediate representation, so the Linalg Dialect can provide an intermediate representation that is easier to optimize without chip information. For example, the Triton intermediate representation can be converted to the Triton dialect (Dialect) based on the MLIR framework, and then based on the triton-shared-opt tool (Triton shared optimization tool), the Triton Dialect can be converted to the corresponding Linalg Dialect through the triton-to-linalg conversion pass (i.e., the conversion optimization process from Triton to Linalg). The Triton dialect is the dialect used to represent computational logic in the MLIR framework. In practice, other conversion and optimization processes can be used to convert the Triton intermediate representation to the Linalg dialect. This is not a limitation here. For example, a direct conversion from the Triton intermediate representation to the Linalg dialect can be performed using a corresponding conversion and optimization process.

[0050] Among them, through optimization based on the Linalg Dialect that does not involve chip information, the performance of operators can be improved. For example, loop fusion can be performed to merge multiple loops into a single loop, reducing the overhead of loop control logic while improving cache utilization. For another example, common subexpression elimination (CSE) can be performed to identify and eliminate duplicate subexpressions, reducing unnecessary calculations. For another example, dead code elimination (DCE) can be performed to remove unused variables, operations, or code snippets, reducing code size and unnecessary calculations. For another example, constant propagation can be performed to propagate constant values ​​into expressions, reducing runtime calculations. For another example, operation fusion can be performed to merge multiple operations into a single operation, reducing the storage and computational overhead of intermediate results. For another example, memory access pattern optimization can be performed to optimize memory access patterns and reduce cache misses. For another example, instruction reordering can be performed to optimize the execution order of instructions and reduce dependency delays. For another example, layout normalization can be performed to convert the data layout into a standard form to facilitate subsequent optimization. For another example, element-wise operation optimization can be performed to optimize element-wise operations by merging operations or eliminating redundant calculations. For another example, matrix operation optimization can be performed to reduce the amount of calculation for matrix operations by adjusting the operation order or layout.

[0051] By optimizing operators based on the Triton intermediate representation without chip information, the operator's performance can be optimized and improved, thereby increasing the operator's operating efficiency on the target chip and boosting computing power. Furthermore, because the optimization of operators based on the Triton intermediate representation is a general optimization that does not involve chip information, the process can be open sourced and does not disclose the target chip's chip information. Therefore, the resources of the open source community can be leveraged to further optimize and expand the optimization process of the Triton intermediate representation-based general optimization of operators without chip information, further improving the performance of compiled operators quickly and effectively.

[0052] S104. Compile the optimized intermediate representation using the Triton compiler corresponding to the target chip to obtain the target code.

[0053] Among them, the Triton compiler corresponding to the target chip can be called through the corresponding compilation interface provided by the manufacturer of the target chip (i.e., the compilation interface corresponding to the target chip), so that the Triton compiler corresponding to the target chip can be used to compile and process the optimized intermediate representation obtained above, thereby obtaining the target code. The target code can be a machine code that can be executed by the target chip. Of course, in some other possible implementations, the Triton compiler corresponding to the target chip can also be pre-configured, so that the Triton compiler corresponding to the target chip can be directly used to compile and process the optimized intermediate representation to obtain the target code during compilation, thereby eliminating the need to call the Triton compiler through the compilation interface, thereby improving compilation efficiency.

[0054] Because chip manufacturers usually hide chip information in the Triton compiler they provide in order to protect the architecture information of their chips, in this embodiment of the application, by calling the Triton compiler corresponding to the target chip, the aforementioned optimized intermediate representation is further optimized in a manner involving chip information, and the operator is finally compiled to obtain the target code. This can optimize the operator's chip information while protecting the manufacturer's chip information, thereby further improving the operator's performance.

[0055] In an embodiment of the present application, for the intermediate representation obtained above that is optimized without involving chip information, the subsequent compilation process based on the Triton compiler can be configured by the manufacturer of the target chip, and is not limited here. For example, the optimized TTGIR can be further optimized with respect to chip information (or hardware-related optimization) by the Triton compiler, and then the optimized TTGIR is converted into a low-level virtual machine (LLVM) IR, and then the LLVM IR is converted into the target code. For another example, the optimized Linalg dialect can be further downgraded and optimized by the Triton compiler, and finally a hardware-specific dialect for hardware is obtained, and then the hardware-specific dialect is optimized with respect to chip information (or hardware-related optimization), and then the optimized hardware-specific dialect is converted into LLVM IR, and then the LLVM IR is converted into the target code. Of course, the above is only a simple example of the process of compiling the optimized intermediate representation based on the Triton compiler to obtain the target code. In actual applications, it can be configured according to actual needs and relevant technologies, and is not limited here.

[0056] Based on the above implementation, the operator compilation method provided by this application can convert the operator source code based on Triton DSL into target code in stages, and optimize the operator source code accordingly to improve the performance of the operator. For example, Figure 2 A flowchart of another operator compilation method provided in an embodiment of the present application.

[0057] like Figure 2 As shown, the compilation process of the operator can be divided into three stages. Among them, in the first stage, the Triton intermediate representation obtained by converting the operator source code based on Triton DSL can be converted into TTGIR or Linalg dialect according to the type of target chip. Specifically, when the target chip is GPGPU, the Triton intermediate representation is converted into TTGIR, and when the target chip is ASIC, the Triton intermediate representation is converted into Linalg dialect. In the second stage, general optimization that does not involve chip information can be performed based on TTGIR or Linalg dialect. In the third stage, the optimized intermediate representation (such as optimized TTGIR or optimized Linalg dialect) can be compiled proprietary by the Triton compiler corresponding to the target chip to obtain the target code. For example, the first and second stages can be implemented in open source, so as to continuously optimize and expand the general optimization process of the operator that does not involve chip information based on the resources of the open source community, thereby improving the performance of the operator.

[0058] The compilation method of the operator provided in the embodiment of the present application divides the compilation process of the operator source code into a general optimization (i.e., optimization that does not involve chip information) part based on the Triton intermediate representation, and a closed-source proprietary optimization part that introduces chip information based on the Triton compiler corresponding to the target chip. Since the general optimization part based on the Triton intermediate representation does not involve chip information, this part of the compilation process can be open source, so as to facilitate optimization and expansion of this part of the compilation process based on the open source community. The closed-source proprietary optimization part of the Triton compiler corresponding to the target chip, its Triton compiler can be provided by the manufacturer of the target chip, and the manufacturer can still adopt a closed-source approach to avoid leakage of its chip information. The compilation method of the operator can realize open source of some stages in the operator compilation process without leaking chip information, so as to optimize and expand the compilation process of the operator based on the open source community, thereby improving the compilation performance of the operator.

[0059] It is obvious that those skilled in the art can make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of the present application.

[0060] It is understandable that, in order to implement the above functions, the electronic device may include hardware structures and / or software modules that perform the corresponding functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0061] The embodiment of the present application can divide the functional modules according to the above method example, so as to obtain a device capable of implementing the above method example. For example, Figure 3 FIG. 1 shows a schematic diagram of a compiling device for an operator provided in an embodiment of the present application. Figure 3 As shown, the device may include: an acquisition module 301, used to obtain the operator source code to be compiled, which is a code written in a domain-specific language based on Triton; a conversion module 302, used to convert the operator source code into a Triton intermediate representation, which is used to represent the calculation logic of the operator source code; an optimization module 303, used to optimize the operator based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation; a compilation module 304, used to compile and process the optimized intermediate representation through the Triton compiler corresponding to the target chip to obtain the target code.

[0062] In one possible implementation, the target chip is a general-purpose graphics processor, and the optimization module 303 is specifically used to convert the Triton intermediate representation into TTGIR, which is an intermediate representation based on the MLIR framework and is used to express the computational logic of the operator source code on the graphics processor; based on TTGIR, the operator is optimized without involving chip information to obtain an optimized intermediate representation.

[0063] In another possible implementation, the target chip is an application-specific integrated circuit, and the optimization module 303 is specifically used to convert the Triton intermediate representation into a linear algebra dialect, which is an intermediate representation used to represent linear algebraic calculation operations in the MLIR framework; based on the linear algebra dialect, the operator is optimized without involving chip information to obtain an optimized intermediate representation.

[0064] In another possible implementation, the optimization module 303 is specifically configured to convert the Triton intermediate representation into a Triton dialect, which is a dialect used to represent computational logic in the MLIR framework; and convert the Triton dialect into a linear algebra dialect.

[0065] In another possible implementation, the conversion module 302 is specifically configured to parse the operator source code into a Python abstract syntax tree; traverse the abstract syntax tree to obtain information nodes in the operator source code; and convert the operator source code into a Triton intermediate representation based on the information nodes.

[0066] In another possible implementation, the chip information includes at least one of the chip's architecture type, the number of cores of the chip, the chip's instruction set extension architecture information, the chip's memory hierarchy information, the chip's memory type, the chip's bus architecture information, and the chip's interface information.

[0067] In another possible implementation, the compilation module 304 is further configured to call the Triton compiler through a compilation interface corresponding to the target chip.

[0068] The embodiment of the present application also provides an electronic device, Figure 4 FIG. 1 shows a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown, the electronic device includes: a bus 401, a processor 402, a memory 403, and a communication interface 404. The processor 402, the memory 403, and the communication interface 404 communicate with each other via the bus 401. The memory 403 stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor 402, the electronic device executes the method provided in the above embodiment. The electronic device can be a server or a terminal device. It should be understood that this application does not limit the number of processors 402 and memories 403 in the electronic device.

[0069] The bus 401 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus 401 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus 401 may include a path for transmitting information between various components of the electronic device (eg, memory 403, processor 402, communication interface 404).

[0070] The processor 402 may include any one or more processors such as a central processing unit, a graphics processing unit, a microprocessor (MP), or a digital signal processor (DSP).

[0071] The memory 403 may include a volatile memory, such as a random access memory (RAM). The memory 403 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0072] The communication interface 404 uses a command distribution module such as, but not limited to, a network interface card, a transceiver, etc. to implement communication between the electronic device and other devices or a communication network.

[0073] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the method provided in the aforementioned embodiment.

[0074] In addition, in the embodiments of the present application, the functional units or modules of the apparatus for implementing the aforementioned method examples may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0075] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0076] Therefore, an embodiment of the present application further provides a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, the processor executes the method provided in the aforementioned embodiment.

[0077] An embodiment of the present application further provides a chip system comprising at least one processor and at least one interface circuit. The processor and the interface circuit may be interconnected via a circuit. For example, the interface circuit may be used to receive signals from another device (e.g., a memory of an electronic device). For another example, the interface circuit may be used to send signals to another device (e.g., a processor).

[0078] For example, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the chip system can perform the various steps in the above embodiment. Of course, the chip system can also include other discrete components, which are not specifically limited in the embodiments of the present application.

[0079] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for compiling an operator, characterized in that: include: Obtain the operator source code to be compiled, where the operator source code is written in a domain-specific language based on Triton; Converting the operator source code into a Triton intermediate representation, where the Triton intermediate representation is used to represent the computational logic of the operator source code; Optimizing the operator without involving chip information based on the Triton intermediate representation to obtain an optimized intermediate representation; The optimized intermediate representation is compiled and processed by the Triton compiler corresponding to the target chip to obtain the target code.

2. The method according to claim 1, characterized in that The target chip is a general-purpose graphics processor, and the optimization of the operator based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation includes: Convert the Triton intermediate representation into a Triton graphics processor intermediate representation TTGIR, where TTGIR is an intermediate representation based on a multi-level intermediate representation (MLIR) framework and is used to express the computational logic of the operator source code on a graphics processor. The operator is optimized based on the TTGIR without involving chip information to obtain the optimized intermediate representation.

3. The method according to claim 1, characterized in that The target chip is an application specific integrated circuit, and the optimization of the operator based on the Triton intermediate representation without involving chip information to obtain an optimized intermediate representation includes: Converting the Triton intermediate representation into a linear algebra dialect, where the linear algebra dialect is an intermediate representation used to represent linear algebra operations in the MLIR framework; The operator is optimized based on the linear algebra dialect without involving chip information to obtain the optimized intermediate representation.

4. The method according to claim 3, characterized in that Converting the Triton intermediate representation into a linear algebra dialect includes: Converting the Triton intermediate representation into a Triton dialect, where the Triton dialect is a dialect used to represent computational logic in the MLIR framework; Converts the Triton dialect to the linear algebra dialect.

5. The method according to any one of claims 1 to 4, characterized in that The converting of the operator source code into Triton intermediate representation includes: Parsing the operator source code into a Python abstract syntax tree; Traversing the abstract syntax tree to obtain information nodes in the operator source code; Based on the information node, the operator source code is converted into the Triton intermediate representation.

6. The method according to claim 1, wherein The chip information includes at least one of the chip's architecture type, the number of cores of the chip, the chip's instruction set extension architecture information, the chip's memory hierarchy information, the chip's memory type, the chip's bus architecture information, and the chip's interface information.

7. The method according to claim 1, characterized in that Before compiling and processing the optimized intermediate representation by the Triton compiler corresponding to the target chip to obtain the target code, the method further includes: The Triton compiler is called through the compilation interface corresponding to the target chip.

8. A compilation device for an operator, characterized in that: include: An acquisition module is used to obtain the operator source code to be compiled. The operator source code is a code written in a domain-specific language based on Triton; A conversion module, configured to convert the operator source code into a Triton intermediate representation, wherein the Triton intermediate representation is used to represent the computational logic of the operator source code; An optimization module, configured to optimize the operator without involving chip information based on the Triton intermediate representation to obtain an optimized intermediate representation; The compilation module is used to compile and process the optimized intermediate representation through the Triton compiler corresponding to the target chip to obtain the target code.

9. An electronic device, characterized in that: The electronic device includes: a processor, a memory and a communication interface; the memory and the communication interface are coupled to the processor, the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Task execution method based on domain-specific language and software development tool chain

    CN119311253A

Cited By

  • Code generation method, electronic equipment and storage medium

    CN121070322A