Compilation method, compilation device, electronic device, and storage medium

Through the multi-level intermediate representation definition method, the problems of high conversion overhead and high development cost of existing compilation frameworks in heterogeneous on-chip systems are solved, an efficient compilation process is achieved, and the computing power utilization of the processing unit is maximized.

CN114461221BActive Publication Date: 2025-09-05BEIJING ESWIN COMPUTING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210102568.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-09-05
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

Existing compilation frameworks have problems such as large conversion overhead, difficulty in collaborative optimization, and high development costs when converting computational graphs into execution instructions for heterogeneous on-chip systems. Especially when the processing units are quite different, the computing power of the processing units cannot be maximized.

Method used

A multi-level intermediate representation definition method is adopted. Through operator fusion and step-by-step addition of hardware platform feature information, the graph-based intermediate representation is gradually optimized into a low-level intermediate representation related to the hardware platform, reducing the conversion span and lowering development costs.

Benefits of technology

It improves compilation efficiency, reduces the conversion overhead from computational graphs to hardware execution instructions, reduces the cost of building domain-specific compilers, and maximizes the computing power utilization of processing units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114461221B_ABST
    Figure CN114461221B_ABST
Patent Text Reader

Abstract

A compilation method, compilation device, electronic device, and storage medium. The compilation method includes: obtaining a first intermediate representation corresponding to the object to be compiled, wherein the first intermediate representation is a graph-based intermediate representation; performing multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation; and obtaining executable instructions corresponding to each processing unit based on the second intermediate representation, wherein the multi-level conversion and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to multiple processing units. In response to the diversity of hardware platforms, the compilation method gradually optimizes and converts the graph-based intermediate representation into a low-level intermediate representation related to the hardware platform, reducing the conversion span from the graph-based first intermediate representation to the executable instructions supported by the processing unit, reducing the conversion overhead from the computational graph to the hardware execution instructions, and reducing the cost of building a compiler for a specific domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a compilation method, a compilation device, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] Deep learning algorithms have been widely used in various industries, including security, smart home, and smart education. The emergence of new neural network architectures and the rapid iteration and update of deep learning algorithms require chips to provide the highest level of computing power. With the continuous improvement of integrated circuit technology, system-on-chip (SoC) design has gradually developed since the mid-1990s. Chip designers can integrate multiple complex functions onto a single silicon wafer, significantly improving system computing power and reducing system power consumption. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a compilation method applicable to heterogeneous devices, wherein the heterogeneous devices include multiple different types of processing units, and the compilation method includes: obtaining a first intermediate representation corresponding to an object to be compiled, wherein the first intermediate representation is a graph-based intermediate representation; performing multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation; and obtaining executable instructions corresponding to each processing unit based on the second intermediate representation; wherein the multi-level conversion and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to the multiple processing units.

[0004] For example, in the compilation method provided in at least one embodiment of the present disclosure, the first intermediate representation includes multiple operation operators, and the first intermediate representation is converted and optimized at multiple levels to obtain a second intermediate representation, including: obtaining at least one reference operator, wherein each reference operator can be directly supported by at least one processing unit; converting the multiple operation operators according to the at least one reference operator to obtain an operator fusion intermediate representation; performing at least one level of conversion and optimization on the operator fusion intermediate representation according to the hardware platform characteristics of each processing unit, and adding hardware platform characteristic information related to the multiple processing units step by step to obtain the second intermediate representation.

[0005] For example, in the compilation method provided in at least one embodiment of the present disclosure, the multiple operation operators are converted according to the at least one reference operator to obtain an operator fusion intermediate representation, including: determining at least one operation operator from the multiple operation operators that executes the function of the target reference operator, wherein the target reference operator is any reference operator among the at least one reference operator; converting the at least one operation operator into an expression form of the target reference operator to obtain the operator fusion intermediate representation.

[0006] For example, in the compilation method provided in at least one embodiment of the present disclosure, converting the at least one operation operator into the expression form of the target reference operator to obtain the operator fusion intermediate representation includes: determining a first namespace; converting the at least one operation operator into the expression form of the target reference operator in the first namespace to obtain the operator fusion intermediate representation.

[0007] For example, in the compilation method provided in at least one embodiment of the present disclosure, determining the first namespace includes: obtaining the first namespace based on a multi-level intermediate representation architecture; and defining the at least one reference operator in the first namespace.

[0008] For example, in the compilation method provided in at least one embodiment of the present disclosure, defining the at least one reference operator in the first namespace includes: defining a description of each reference operator, operation parameters and parameter types of each reference operator, and operation results and result types of each reference operator in the first namespace.

[0009] For example, in the compilation method provided by at least one embodiment of the present disclosure, the at least one reference operator includes an operator not supported by the first intermediate representation.

[0010] For example, in the compilation method provided in at least one embodiment of the present disclosure, the operator fusion intermediate representation is converted and optimized at least once according to the hardware platform characteristics of each processing unit, and the hardware platform characteristic information related to the multiple processing units is added step by step to obtain the second intermediate representation, including: adding a description directly related to the hardware platform characteristics of each processing unit to the operator fusion intermediate representation through the at least once conversion and optimization to obtain the second intermediate representation, wherein the second intermediate representation contains the hardware platform characteristic information of each processing unit.

[0011] For example, in the compilation method provided in at least one embodiment of the present disclosure, through the at least one level of conversion and optimization, a description directly related to the hardware platform characteristics of each processing unit is added to the operator fusion intermediate representation to obtain the second intermediate representation, including: determining at least one second namespace corresponding one-to-one to the at least one level of conversion and optimization; converting the multiple first intermediate operators included in the operator fusion intermediate representation into expressions in the corresponding second namespace step by step, wherein the expression of each first intermediate operator in the corresponding second namespace contains the hardware platform characteristic information of the processing unit that executes each first intermediate operator; and using the intermediate representation obtained by the last level of conversion and optimization as the second intermediate representation.

[0012] For example, in the compilation method provided in at least one embodiment of the present disclosure, determining at least one second namespace corresponding one-to-one to the at least one level of conversion and optimization includes: obtaining the at least one second namespace based on a multi-level intermediate representation architecture; defining expression forms corresponding to the multiple first intermediate operators in each second namespace, wherein the expression form corresponding to each first intermediate operator in each second namespace adds hardware platform characteristic information of the processing unit that executes the first intermediate operator.

[0013] For example, in the compilation method provided in at least one embodiment of the present disclosure, the expression forms corresponding to the multiple first intermediate operators are defined in each second namespace, including: for each first intermediate operator, determining the processing unit that executes the first intermediate operator; adding an execution unit identifier for the first intermediate operator according to the processing unit; determining the hardware platform characteristic information of the processing unit, wherein the hardware platform characteristic information includes hardware operation information or hardware optimization information, the hardware operation information includes one or more of the data shape, data precision, and data storage location corresponding to the processing unit, and the hardware optimization information includes one or more combinations of hardware storage analysis, hardware power consumption minimization, data blocking, and hardware storage quantization; according to the hardware platform characteristic information of the processing unit, the expression form corresponding to the first intermediate operator is defined in each second namespace, wherein different hardware platform characteristic information is defined in different second namespaces.

[0014] For example, in the compilation method provided in at least one embodiment of the present disclosure, the second intermediate representation includes multiple second intermediate operators, and according to the second intermediate representation, the executable instructions corresponding to each processing unit are obtained, including: determining the processing unit corresponding to each second intermediate operator according to the second intermediate representation, wherein each second intermediate operator is executed on the corresponding processing unit, and the second intermediate representation includes an execution unit identifier corresponding to each second intermediate operator; according to the instruction set type and rules of each processing unit, extracting the hardware platform characteristic information of at least one second intermediate operator executed on each processing unit from the second intermediate representation to generate executable instructions corresponding to each processing unit.

[0015] For example, in the compilation method provided in at least one embodiment of the present disclosure, the executable instructions corresponding to each processing unit are obtained according to the second intermediate representation, including: converting the second intermediate representation into an underlying virtual intermediate representation; and generating executable instructions corresponding to each processing unit according to the underlying virtual intermediate representation.

[0016] For example, in the compilation method provided in at least one embodiment of the present disclosure, obtaining a first intermediate representation corresponding to the object to be compiled includes: obtaining operation information of the object to be compiled, wherein the operation information is in the form of a computational graph of the object to be compiled; and converting the object to be compiled into the first intermediate representation according to the operation information.

[0017] For example, in the compilation method provided in at least one embodiment of the present disclosure, the object to be compiled is converted into the first intermediate representation according to the operation information, including: compiling the operation information through an accelerated linear algebra compiler to obtain a high-order optimized intermediate representation; and using the high-order optimized intermediate representation as the first intermediate representation.

[0018] For example, in the compilation method provided in at least one embodiment of the present disclosure, the object to be compiled is used to perform neural network calculations.

[0019] For example, in the compilation method provided in at least one embodiment of the present disclosure, the multiple different types of processing units include at least two of a central processing unit, a graphics processing unit, a neural network processing unit, and a digital processing unit.

[0020] At least one embodiment of the present disclosure provides a compilation device applicable to a heterogeneous device, wherein the heterogeneous device includes multiple different types of processing units, and the compilation device includes: an acquisition unit, configured to obtain a first intermediate representation corresponding to a to-be-compiled object, wherein the first intermediate representation is a graph-based intermediate representation; a conversion unit, configured to perform multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation; and a generation unit, configured to obtain executable instructions corresponding to each processing unit based on the second intermediate representation, wherein the multi-level conversion and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to the multiple processing units.

[0021] At least one embodiment of the present disclosure provides an electronic device, comprising: a memory, which non-transitorily stores computer-executable instructions; and a processor, configured to execute the computer-executable instructions, wherein the computer-executable instructions, when executed by the processor, implement the compilation method according to at least one embodiment of the present disclosure.

[0022] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the compilation method according to at least one embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.

[0024] Figure 1 is a schematic architectural diagram of a heterogeneous system on a chip;

[0025] Figure 2 Compilation flow chart for the computation graph;

[0026] Figure 3 A schematic flowchart of a compilation method provided in at least one embodiment of the present disclosure;

[0027] Figure 4 A schematic flowchart of step S20 provided in at least one embodiment of the present disclosure;

[0028] Figure 5 A processing flow chart of a compilation method provided in one embodiment of the present disclosure;

[0029] Figure 6A A schematic diagram of a first intermediate representation provided for an embodiment of the present disclosure;

[0030] Figure 6B A schematic diagram describing a first namespace provided in an embodiment of the present disclosure;

[0031] Figure 6C A schematic diagram describing a reference operator provided in one embodiment of the present disclosure;

[0032] Figure 6D A schematic diagram of an intermediate representation of operator fusion provided in one embodiment of the present disclosure;

[0033] Figure 7A A schematic diagram describing a second namespace provided in an embodiment of the present disclosure;

[0034] Figure 7B A schematic diagram describing a second namespace provided in an embodiment of the present disclosure;

[0035] Figure 7C A schematic diagram of a second intermediate representation provided in accordance with an embodiment of the present disclosure;

[0036] Figure 8 A schematic block diagram of a compiling device provided in accordance with at least one embodiment of the present disclosure;

[0037] Figure 9 A schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure;

[0038] Figure 10A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0040] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by persons of ordinary skill in the field to which this disclosure belongs. The words "first", "second" and similar terms used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0041] In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of some known functions and components.

[0042] Due to the limitations of Moore's Law, the conventional method of increasing processor frequency to improve chip computing power has little effect. As computing becomes more diversified, more and more scenarios are beginning to introduce a variety of different processing units such as CPU (Central Processing Unit), DSP (Digital Signal Processing Unit), GPU (Graphics Processing Unit), ASIC (Application Specific Integrated Circuit), and FPGA (Field Programmable Gate Array) to accelerate computing, thus giving rise to heterogeneous computing.

[0043] Heterogeneous computing on a chip (SoC) involves combining chips with different manufacturing processes and architectures to solve computing problems. Technically, this is typically accomplished using a heterogeneous multi-core architecture to build a SoC, boosting the overall computing power. Heterogeneous SoCs leverage collaborative computing and mutual acceleration between processing units with different instruction sets and architectures, thus overcoming the bottlenecks inherent in the development of individual processing units.

[0044] Figure 1 The following is a schematic diagram of the architecture of a heterogeneous system-on-chip (SoC). For example, this SoC can be used in complex deep learning applications.

[0045] like Figure 1 As shown, a possible heterogeneous system-on-chip 100 consists of a central processing unit 110, a neural network processing unit 120, a digital signal processing unit 130, a first memory 140, a second memory 150 and necessary communication links 160.

[0046] Computers, mobile phones, embedded devices, etc. generally have heterogeneous architectures, including at least one central processing unit 110 and other heterogeneous processing units. For example, these devices are usually hosted by a central processing unit 110, which transmits specific computing tasks to other processing units; the digital signal processing unit 130 has the characteristics of fast processing speed, high flexibility and strong specialization, and has more advantages than the central processing unit 110 in high-speed computing scenarios and digital processing; the neural network processing unit 120 imitates human neurons and synapses at the circuit layer, and uses a deep learning instruction set to directly process large-scale neurons and synapses. One instruction completes the processing of a group of neurons. Compared with the central processing unit 110 and the digital signal processing unit 130, the neural network processing unit 120 realizes the integration of storage and computing through synaptic weights, thereby improving operating efficiency, and is usually used to execute related instructions of deep learning; the first memory 140 is used to store intermediate calculation results and serve as a cache for the second memory 150, generally having a faster access rate than the second memory 150; the communication link 160 is used to transmit data and control commands between the central processing unit 110, the neural network processing unit 120, the digital signal processing unit 130, the first memory 140 and the second memory 150.

[0047] In addition, the central processing unit 110, the neural network processing unit 120, and the digital signal processing unit 130 may also have internal storage areas. Generally, the access rate of the internal storage areas is better than the access rate of the first memory 140, and the access rate of the first memory 140 is better than the access rate of the second memory 150.

[0048] Operations in neural network models, such as weighted summation, convolution, and activation functions, are organized into a graph structure consisting of nodes and edges, commonly called a computation graph. A computation graph is a directed graph in which nodes represent specific operations or variables in the neural network, and edges represent the flow of data. The computation graph graphically represents the computational process. For example, variables can provide their values ​​to operations, and operations can output computational results and provide them to other operations. In this way, each node in the graph defines a function of the variable, also known as an operator. The values ​​entering and exiting a node are called tensors. Tensors have different dimensions and include scalars, vectors, matrices, and higher-level tensors. For example, a 0-dimensional tensor is also called a scalar.

[0049] Figure 2 Compilation flowchart for the computation graph.

[0050] like Figure 2 As shown, converting the computation graph 210 into execution instructions in the heterogeneous system-on-chip 100 requires processing by the compiler 200 .

[0051] like Figure 2 As shown, first, in step 210, the compiler 200 constructs an intermediate representation based on the computation graph; then, in step 220, the compiler 200 optimizes the constructed intermediate representation to obtain an executable file, which contains executable instructions on the target hardware platform (such as the heterogeneous system on chip 100), or is called an executable instruction sequence.

[0052] For example, an intermediate representation (IR) refers to a data structure used for analysis and conversion between the time a compiler receives a source program and performs semantic analysis on it and the time it outputs an executable program.

[0053] The heterogeneous system-on-chip 100 is composed of a variety of different processing units. How to design a unified compilation framework to support different hardware back-end platforms is a huge challenge brought to the design of deep learning compilers by the rapid development of deep learning algorithms and hardware.

[0054] Existing compilation frameworks attempt to address the universality issue. For example, Google's TensorFlow platform uses XLA (Accelerated Linear Algebra) to compile computational graphs into a graph-based HLO (High Level Optimizer) IR, similar to a high-level language. HLO IR is then optimized to generate LLVM (Low Level Virtual Machine) IR, a low-level language similar to a RISC instruction set. LLVM IR is then compiled into assembly language for various hardware platforms (i.e., the different processing units of heterogeneous systems-on-chip).

[0055] However, this compilation method has the following three problems:

[0056] 1. HLO IR is an intermediate representation based on the computation graph. Therefore, this intermediate representation differs significantly from the instructions supported by the hardware in terms of structure and abstraction level. The conversion overhead from HLO IR to LLVM IR is relatively high, and co-optimization is difficult.

[0057] 2. HLO IR defines universal atomic-granularity operators, such as some operators that perform basic calculations, such as addition, subtraction, dot multiplication, broadcasting, maximum value, etc. This is also a major feature of HLO IR. This method can easily represent the calling relationship between computational graphs, and implement complex calculations based on the characteristics of different hardware through the combination of basic operators supported by HLO. However, when used in heterogeneous systems-on-chip for neural network computing, some processing units can directly support complex operators or operators that HLO IR does not yet support. The atomic-granularity operators defined by HLO IR cannot maximize the computing power of the processing units, increasing the conversion overhead.

[0058] 3. Heterogeneous SoCs are composed of different processing units, which require different forms of intermediate representations. For example, HLO IR needs to be converted into LLVM IR adapted to different processing units. This conversion is inefficient and increases the compiler development cost.

[0059] At least one embodiment of the present disclosure provides a compilation method, compilation device, electronic device, and non-transitory computer-readable storage medium. The compilation method includes: obtaining a first intermediate representation corresponding to an object to be compiled, wherein the first intermediate representation is a graph-based intermediate representation; performing multi-level transformation and optimization on the first intermediate representation to obtain a second intermediate representation; and obtaining executable instructions corresponding to each processing unit based on the second intermediate representation, wherein the multi-level transformation and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to the various processing units.

[0060] This compilation method provides a unified multi-level intermediate representation definition method. In response to the diversity of hardware platforms, it gradually optimizes and converts the graph-based intermediate representation into a low-level intermediate representation related to the hardware platform, reducing the conversion span between the first graph-based intermediate representation and the executable instructions supported by the processing unit, reducing the conversion overhead from the computational graph to the hardware execution instructions, and reducing the cost of building a compiler for a specific domain.

[0061] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0062] Figure 3 A schematic flowchart of a compilation method provided in at least one embodiment of the present disclosure.

[0063] For example, the compilation method is applicable to a heterogeneous device, for example, a heterogeneous system-on-chip, and for example, the heterogeneous device may include multiple different types of processing units. For example, the multiple different types of processing units include at least two of a central processing unit, a graphics processing unit, a neural network processing unit, and a digital processing unit.

[0064] For example, one possible architecture for heterogeneous devices is Figure 1 The architecture shown, but it should be noted that Figure 1 The heterogeneous system-on-chip shown is one possible architecture, and this disclosure does not limit it. In this disclosure, a heterogeneous device is composed of multiple different types of processing units. Different types of processing units refer to processing units with different instruction sets or architectures. For example, a heterogeneous device can also adopt an architecture of central processing unit + graphics processing unit, or a central processing unit + neural network processing unit + digital signal processing unit, etc. Each architecture can have more than one processing unit of each type, and this disclosure does not specifically limit the type or number of processing units.

[0065] For example, Figure 3 As shown, the compilation method provided by the embodiment of the present disclosure includes steps S10 to S30.

[0066] In step S10, a first intermediate representation corresponding to the object to be compiled is obtained.

[0067] For example, the first intermediate representation is a graph-based intermediate representation.

[0068] For example, step S10 may include: obtaining operation information of the object to be compiled, where the operation information is in the form of a computation graph of the object to be compiled; and converting the object to be compiled into a first intermediate representation according to the operation information.

[0069] For example, the object to be compiled is used to perform neural network calculations. For example, the object to be compiled can be an algorithm program for performing neural network calculations, and the algorithm program is written in any feasible high-level language, such as Python.

[0070] For example, according to the neural network operations to be performed by the compiled object, such as weighted summation, convolution, activation function, etc., they are arranged into a graph structure composed of nodes and edges. This graph structure is the operation information in the form of a computational graph.

[0071] For example, an existing compilation framework can be used to process the computation information to generate an intermediate representation. For example, converting the object to be compiled into a first intermediate representation based on the computation information can include: compiling the computation information using an accelerated linear algebra compiler to obtain a high-order optimized intermediate representation; and using the high-order optimized intermediate representation as the first intermediate representation.

[0072] For example, the operation information is compiled into HLO IR through XLA, and HLO IR is used as the first intermediate representation here.

[0073] It should be noted that those skilled in the art may also adopt other compilation frameworks and compilation methods to obtain the first intermediate representation here. The obtained first intermediate representation may be a graph-based intermediate representation, or in other words, the first intermediate representation is a higher-order intermediate representation that is closer to a high-level language.

[0074] Afterwards, on this basis, targeting the diversity of hardware platforms, the first intermediate representation is gradually optimized and converted into intermediate representations related to the hardware platform.

[0075] In step S20, the first intermediate representation is subjected to multi-level transformation and optimization to obtain a second intermediate representation.

[0076] For example, the multi-level transformation and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to various processing units.

[0077] For example, the input of each level of transformation and optimization is an intermediate representation, and the output is also an intermediate representation. The output intermediate representation performs operator fusion relative to the input intermediate representation, or adds some hardware platform feature information.

[0078] For example, when the first intermediate representation is a graph-based intermediate representation, the first intermediate representation includes multiple operation operators, and the multiple operation operators are atomic-granularity operators, such as operators that implement basic operations such as addition, subtraction, dot multiplication, and maximum value.

[0079] The operators directly supported by the hardware processing unit may be operation operators, or they may directly support more complex operators composed of multiple operation operators, such as the activation function operator directly supported by the neural network processing unit. In this case, it is necessary to perform operator fusion on the operation operators, combining one or more operation operators that implement complex operator functions into an operator form directly supported by the processing unit, so as to maximize the computing power of each processing unit in the heterogeneous device and reduce the conversion span between the graph-based intermediate representation and the lower-level intermediate representation.

[0080] In addition to operator fusion, the present disclosure also adds hardware platform feature information related to various processing units through at least one level of conversion and optimization. For example, this hardware platform feature information provides the hardware operation information required by the processing unit to run each operator, such as the shape of input and output data, data storage location, data accuracy, etc. In addition, this hardware platform feature information can also provide hardware optimization information, including conventional optimizations such as dead code elimination, constant folding, and common subexpression elimination. It can also perform more specific custom optimizations based on the specific needs of each processing unit and the hardware characteristics of each processing unit.

[0081] For example, multi-level optimization and conversion may include two or more levels of conversion and optimization. For example, the first intermediate representation is converted and optimized at N levels, where N is a positive integer greater than or equal to 2. In the first-level conversion, some optimizations and operator fusions that are common to each processing unit are performed. In the subsequent N-1-level conversion and optimization, the hardware characteristics are combined to specifically add the required hardware platform characteristic information. Thus, the first-level conversion and optimization are decoupled from the hardware characteristics, focusing on the conversion and optimization of operator combination and operator definition. In the subsequent N-1-level conversion, descriptions directly related to the hardware characteristics are gradually added. This method decouples the computational graph from the hardware platform environment of the final execution, avoiding the completion of both operator-level work and work strongly related to hardware in the first-level conversion, improving the scalability of the compilation method, and reducing the cost of building compilers for specific fields.

[0082] For example, Figure 4 This is a schematic flow chart of step S20 provided in at least one embodiment of the present disclosure. Figure 4 As shown, step S20 at least includes steps S201-S203.

[0083] In step S201, at least one reference operator is acquired.

[0084] For example, each reference operator can be directly supported by at least one processing unit.

[0085] For example, statistics can be collected in advance on neural network models of various structures to determine commonly used operators that may be required for calculations of various neural network models, such as activation function operators, matrix convolution operators, pooling operators, etc. Furthermore, from these statistically obtained operators, one or more operators that can be directly supported by at least one processing unit in the heterogeneous device are determined as reference operators. For example, if the neural network processing unit can directly support the operation of the activation function operator, the activation function operator can be used as the reference operator.

[0086] For example, the reference operator may also include operators not supported by the first intermediate representation. For example, the first intermediate representation does not directly support the matrix convolution operator. In the first intermediate representation, the matrix convolution operator is implemented by combining multiple operators. However, if the neural network processing unit can directly support the matrix convolution operation, the reference operator may also include the matrix convolution operator.

[0087] For example, the reference operator needs to be an operator that can be directly supported by at least one processing unit. The reference operator can change according to the changes in the different types of processing units included in the heterogeneous device, that is, different heterogeneous devices may correspond to different reference operators.

[0088] For example, the reference operator in the present disclosure can be an operator with a granularity greater than that of the operation operator, for example, the function of the reference operator is implemented by a combination of multiple operation operators, or the reference operator can also include operators with the same granularity as the operation operator, for example, basic operations such as addition and subtraction directly supported by the processing unit, and the reference operator can also include operators that perform these basic operations.

[0089] In step S202, multiple operation operators are converted according to at least one reference operator to obtain an operator fusion intermediate representation.

[0090] For example, step S202 may include: determining at least one operation operator that performs the function of the target reference operator from multiple operation operators, wherein the target reference operator is any reference operator among the at least one reference operator; converting at least one operation operator into an expression form of the target reference operator to obtain an operator fusion intermediate representation.

[0091] For example, the first intermediate representation includes multiple operators, which can be combined to complete the functions of the activation function operator and the matrix convolution operator. Of course, the first intermediate representation can also include more calculations. Here, the activation function operator and the matrix convolution operator are used as examples for description. For example, the A operators that complete the functions of the activation function operator are converted into the expression form of the activation function operator, and the B operators that complete the functions of the matrix convolution operator are converted into the expression form of the matrix convolution operator. Thus, an operator fusion intermediate representation is obtained, where A and B are positive integers, and the activation function operator and the matrix convolution operator are reference operators.

[0092] For example, converting at least one operation operator into an expression form of a target reference operator to obtain an operator fusion intermediate representation may include: determining a first namespace; converting at least one operation operator into an expression form of a target reference operator in the first namespace to obtain an operator fusion intermediate representation.

[0093] For example, determining the first namespace may include: obtaining the first namespace based on a multi-level intermediate representation architecture; and defining at least one reference operator in the first namespace.

[0094] For example, the Multi-Level Intermediate Representation (MLIR) architecture aims to provide a multi-level IR format, leveraging its modularity and scalability to address interoperability issues between IRs and compile them into assembly language for specific hardware platforms in a highly consistent manner. MLIR is an infrastructure rather than a specific IR or IR definition method. It emphasizes the continuous optimization and transformation of computational graphs through multi-level IRs, ultimately generating a hardware-executable language. Users need to customize the specific IR based on heterogeneous SoCs.

[0095] Here, we use the Dialect provided by MLIR as a base class or template, defined using MLIR's TableGen syntax. Classes in TableGen are similar to classes in C++, allowing Dialect to serve as templates or base classes from which subclasses can be derived. For example, we define the subclass EHLO_Dialect, which derives from Dialect to construct the first namespace. The Dialect template is responsible for defining various operations and analyses, while also being extensible. For example, the first namespace defined by EHLO_Dialect, which derives from Dialect, is called "ehlo."

[0096] It should be noted that the derived subclass is named EHLO_Dialect and the first namespace is called "ehlo", but the present disclosure does not limit the name of the derived subclass and the namespace.

[0097] For example, defining at least one reference operator in the first namespace may include: defining a description of each reference operator, operation parameters and parameter types of each reference operator, and operation results and result types of each reference operator in the first namespace.

[0098] For example, in the first namespace ehlo, basic parameters such as the data type used to describe the edges of the computation graph (i.e., tensor data) are defined. In addition, the main structure of each reference operator is also defined. For example, each reference operator is derived from the Op (operator) class. Op is the main structure for defining operators provided by MLIR. The derived reference operator is specifically described by specifying the operator name "mnemonic" and the operator constraints "traits".

[0099] For example, each reference operator is defined in the first namespace with a comment "summary", a description "description", the required operation parameters and the operation parameter type ("arguments"), the operation result and the operation result type ("results"). For a detailed description of the reference operator and the first namespace, please refer to the following text. Figure 6B-6D The description in , will not be repeated here.

[0100] For example, the entire first intermediate representation is traversed, and all operators are converted to reference operators defined in the first namespace, that is, replaced with expressions of reference operators defined in the first namespace. Thus, the entire first intermediate representation is converted to expressions defined in the first namespace, obtaining an operator-fused intermediate representation.

[0101] For example, in the operator fusion intermediate representation, complex calculations implemented by multiple operators can be represented by a reference operator in the operator fusion intermediate representation; in addition to these operators that combine to complete complex calculations, for some operators directly supported by the hardware, such as the maximum operator directly supported by the neural network processing unit, these operators are also converted into expressions of the corresponding reference operators in the first namespace in the operator fusion intermediate representation.

[0102] In addition, in the first namespace, some optimizations common to various processing units may also be defined, such as constant folding, dead code elimination, common sub-expression elimination, etc.

[0103] The present disclosure provides a general conversion and definition method for operator fusion. In this level of conversion and optimization, operator fusion and optimization that are common to each hardware unit are performed, and small-grained operation operators in the first intermediate representation are combined into complex operators that can be directly supported by the processing unit. It has strong scalability and reduces the conversion span from the graph-based first intermediate representation to the executable instructions supported by the processing unit.

[0104] For example, after the first level conversion and optimization is completed in step S202, the subsequent N-1 level conversion and optimization is completed in step S203, and the required hardware platform characteristic information is added in a targeted manner in combination with the hardware characteristics.

[0105] In step S203, according to the hardware platform characteristics of each processing unit, the operator fusion intermediate representation is converted and optimized at least one level, and the hardware platform characteristic information related to the various processing units is added level by level to obtain a second intermediate representation.

[0106] For example, step S203 may include: through at least one level of conversion and optimization, adding a description directly related to the hardware platform characteristics of each processing unit to the operator fusion intermediate representation to obtain a second intermediate representation, wherein the second intermediate representation contains hardware platform characteristic information of each processing unit.

[0107] For example, at least one level of conversion and optimization primarily targets the hardware characteristics of each processing unit. For example, the number of conversion stages can be designed based on actual conversion and optimization requirements, with each stage performing further hardware optimization and conversion on the intermediate representation output by the previous stage. Multi-level conversion and optimization can reduce the conversion span from the intermediate representation of operator fusion to executable instructions supported by the hardware platform, thereby lowering conversion overhead.

[0108] For example, through at least one level of conversion and optimization, a description directly related to the hardware platform characteristics of each processing unit is added to the operator fusion intermediate representation to obtain a second intermediate representation, which may include: determining at least one second namespace corresponding one-to-one to at least one level of conversion and optimization; converting the multiple first intermediate operators included in the operator fusion intermediate representation step by step into expressions in the corresponding second namespace, wherein the expression of each first intermediate operator in the corresponding second namespace contains the hardware platform characteristic information of the processing unit that executes each first intermediate operator; and using the intermediate representation obtained by the last level of conversion and optimization as the second intermediate representation.

[0109] For example, based on the hardware characteristics of each processing unit, such as the supported data precision, the storage location of the input data and output data during operator operation, the step size information of the convolution operator, the pooling type of the pooling operator (maximum pooling, average pooling), etc., corresponding hardware platform characteristic information is added to each first intermediate operator to give each first intermediate operator a more specific expression.

[0110] For example, the first intermediate operator here is the expression of the aforementioned reference operator in the first namespace.

[0111] For example, determining at least one second namespace corresponding one-to-one to at least one level of transformation and optimization may include: obtaining at least one second namespace based on a multi-level intermediate representation architecture; defining expression forms corresponding to multiple first intermediate operators in each second namespace, wherein the expression form corresponding to each first intermediate operator in each second namespace adds hardware platform characteristic information of the processing unit that executes the first intermediate operator.

[0112] Here, we also use the Dialect provided by MLIR as a base class or template, using MLIR's TableGen syntax for definition, and using Dialect as a template or base class to derive subclasses. For example, we define a subclass called EIR_Dialect, which derives from Dialect to construct a second namespace. The Dialect template is responsible for defining various operations and analyses, while also being extensible. For example, the namespace defined by EIR_Dialect, which derives from Dialect, is called "eir."

[0113] For example, when multiple second namespaces exist, different second namespaces are derived from different derived subclasses, all of which are derived from Dialect. For example, different second namespaces can have different names. For example, different second namespaces can implement different hardware optimizations based on the definition of different hardware platform characteristics.

[0114] It should be noted that the derived subclass here is named EIR_Dialect, and the corresponding second namespace is called "eir", but the present disclosure does not limit the name of the derived subclass and the name of the namespace.

[0115] For example, defining the expression forms corresponding to multiple first intermediate operators in each second namespace may include: determining, for each first intermediate operator, a processing unit that executes the first intermediate operator; adding an execution unit identifier for the first intermediate operator based on the processing unit; determining the hardware platform characteristic information of the processing unit, wherein the hardware platform characteristic information includes hardware operation information or hardware optimization information, the hardware operation information includes one or more of the data shape, data accuracy, and data storage location corresponding to the processing unit, and the hardware optimization information includes one or more combinations of hardware storage analysis, hardware power consumption minimization, data blocking, and hardware storage quantization; based on the hardware platform characteristic information of the processing unit, defining the first intermediate operator and the hardware platform characteristic information of the first intermediate operator in each second namespace, wherein different hardware platform characteristic information is defined in different second namespaces.

[0116] For example, for a first intermediate operator among multiple first intermediate operators, if the first intermediate operator can be executed in multiple processing units, the processing unit that executes the first intermediate operator is analyzed and determined from the perspectives of execution efficiency, resource consumption, resource allocation, etc. For example, if the first intermediate operator is an activation function operator, and the activation function operator can be directly supported by both a graphics processing unit and a neural network processing unit, and the neural network processing unit has higher execution efficiency, then the processing unit corresponding to the activation function operator is determined to be the neural network processing unit, and a corresponding execution unit identifier is added to the activation function operator to mark the activation function operator as being executed by the neural network processing unit.

[0117] For example, after determining the processing unit that executes the first intermediate operator, some hardware operation information required by the processing unit during the operator operation process and hardware optimization information required for specific hardware optimization can be determined.

[0118] For example, hardware operation information provides the hardware parameters required by the processing unit to execute the operator, such as the data shape of the input data and output data, including the dimension and data type of the tensor, etc. In addition, it can also include data precision, data storage location, etc.

[0119] For example, hardware optimization information includes common optimizations such as constant folding, dead code elimination, and common subexpression elimination. It can also include customized optimization information tailored to hardware characteristics, such as hardware memory analysis, hardware memory quantization, and data partitioning. These optimizations can significantly reduce memory transfers and resource waste. Furthermore, hardware optimization information can also include memory limitations and minimized hardware power consumption based on the characteristics of the processing unit.

[0120] Hardware optimization information and hardware operation information are directly related to the processing unit, and different processing units may have different hardware platform characteristic information. For special hardware platforms, such as digital processing units, if special conversions need to be performed on the digital processing units, such as assembling the format of the input and output parameters of the operator according to the protocol, assigning the ID (identification number) of the operator, etc., a new level of conversion can be added to add these hardware platform characteristic information. Similarly, more levels of conversion and optimization can be added according to specific needs. Therefore, the compilation method provided by the present disclosure has better scalability and flexibility. When it is necessary to increase the hardware optimization of the processing unit, a new level of conversion can be added to achieve it, or the corresponding hardware platform characteristic information can be added in the namespace.

[0121] For example, the hardware platform feature information added at each level is different, that is, different hardware optimizations are performed in the conversion and optimization at different levels.

[0122] For example, in the second namespace eir for converting and optimizing the intermediate representation of operator fusion, storage hierarchy attributes for describing the edges of the computational graph (i.e., tensor data) are defined. The storage hierarchy attributes indicate the storage location of the operator data in the heterogeneous device. In addition, the main structure of each intermediate conversion operator is also defined in the second namespace eir. For example, the intermediate conversion operator refers to the operator expression corresponding to each first intermediate operator in the second namespace eir. For example, each intermediate conversion operator is derived from the Op (operator) class. Op is the main structure for defining operators provided by MLIR. The derived intermediate conversion operator is specifically described by specifying the operator name "mnemonic" and the operator constraint "traits".

[0123] For example, a data conversion precision module can be defined in the second namespace eir. This module is used to convert data of one precision type to data of another precision type. This module includes three properties: truncate, scale, and offset. By specifying these three properties during precision conversion, data of a specified precision type can be obtained.

[0124] Similarly, various other modules for implementing hardware optimization can be defined in the second namespace eir to add corresponding hardware optimization information.

[0125] For details about the intermediate conversion operator and the second namespace, please refer to the following text Figures 7A-7C The description in , will not be repeated here.

[0126] For example, the present disclosure provides a universal definition and expression method that can uniformly describe the hardware characteristics of different processing units. Different processing units can be assigned different specific parameters. That is, a universal expression form is defined in the second namespace, and the parameters corresponding to the processing unit can be written during specific instantiation.

[0127] For example, the entire operator fusion intermediate representation is traversed. For each level of transformation and optimization, the intermediate transformation operator (or first intermediate operator) in the intermediate representation output by the previous level of transformation and optimization is converted into an expression defined in the corresponding second namespace, with hardware platform feature information added to this expression. Thus, the operator fusion intermediate representation is gradually converted into a second intermediate representation, which is closer to the underlying expression and each operator in the second intermediate representation is assigned specific hardware parameters.

[0128] For example, in the compilation method provided by at least one embodiment of the present disclosure, the specific implementation of each level of conversion and optimization is achieved through a multi-level intermediate conversion architecture (MLIR). For example, as described above, the derived classes EHLO_Dialect and EIR_Dialect are defined to construct a first namespace and a second namespace. The MLIR Dialect base class places all intermediate representations in the same namespace (e.g., EHLO_Dialect places all intermediate representations in the first namespace mlir::ehlo, and EIR_Dialect places all intermediate representations in the second namespace mlir::eir), and defines corresponding productions for each intermediate representation. That is, EHLO_Dialect defines how a certain operator is converted from the expression of the operation operator in the first intermediate representation to the expression of the reference operator in the operator fusion intermediate representation, while EIR_Dialect defines how a certain operator is converted from the expression of the first intermediate operator in the operator fusion intermediate representation to the expression in the second namespace. The filling of certain information is also completed in this process, or in a subsequent optimization process, thereby completing the conversion and optimization between different levels. During the compilation process, the base class Dialect is used to traverse the operation information (that is, the calculation graph) and complete the conversion and optimization in units of operators.

[0129] The present disclosure provides a general conversion and definition method for hardware optimization. In this at least one level of conversion and optimization, optimizations strongly related to each hardware unit are performed, hardware calculation information required for operator execution is added step by step, and customized hardware optimization is performed. The higher-order first intermediate representation is converted to the underlying hardware step by step, reducing the conversion overhead from computational graphs to hardware execution instructions and reducing the cost of building compilers for specific fields.

[0130] For example, after obtaining the intermediate representation related to the hardware, that is, the second intermediate representation, the compiler generates executable instructions according to the target processing unit corresponding to each second intermediate operator in the second intermediate representation.

[0131] In step S30, executable instructions corresponding to each processing unit are obtained according to the second intermediate representation.

[0132] For example, in some embodiments, the executable instructions corresponding to each processing unit are obtained directly according to the second intermediate representation.

[0133] For example, the second intermediate representation includes multiple second intermediate operators, and step S30 may include: determining the processing unit corresponding to each second intermediate operator based on the second intermediate representation, wherein each second intermediate operator is executed on the corresponding processing unit, and the second intermediate representation includes the execution unit identifier corresponding to each second intermediate operator; according to the instruction set type and rules of each processing unit, extracting the hardware platform characteristic information of at least one second intermediate operator executed on each processing unit from the second intermediate representation to generate executable instructions corresponding to each processing unit.

[0134] For example, in the second intermediate representation, the execution unit identifier corresponding to each second intermediate operator is added through step S20, and the execution unit identifier marks the processing unit corresponding to each second intermediate operator.

[0135] For example, according to the instruction set type and rules of each processing unit, the relevant hardware parameters of one or more second intermediate operators executed on each processing unit are extracted from the second intermediate representation, and the executable instructions corresponding to each processing unit are generated, and finally compiled to obtain a binary executable file targeting heterogeneous devices.

[0136] For example, in other embodiments, some operators need to run on a general-purpose platform such as a CPU. The second intermediate representation can be converted into a low-level virtual intermediate representation (LLVM IR) to enable the second intermediate representation to be connected to the existing LLVM IR facility, and executable instructions can be obtained using an existing compiler.

[0137] For example, step S30 may include: converting the second intermediate representation into an underlying virtual intermediate representation; and generating executable instructions corresponding to each processing unit according to the underlying virtual intermediate representation.

[0138] Figure 5 A processing flow chart of a compilation method provided in one embodiment of the present disclosure.

[0139] like Figure 5As shown, first, operation information 310 in the form of a computation graph is obtained, and the operation information 310 is compiled by an accelerated linear algebra compiler to obtain an HLO IR, which is used as a first intermediate representation 320. For example, the specific process can refer to the description of step S10.

[0140] Afterwards, the first intermediate representation 320 is subjected to a two-level transformation and optimization 330 to obtain a second intermediate representation 340. For example, the two-level transformation and optimization includes a first-level transformation and optimization 331 and a second-level transformation and optimization 332. As mentioned above, the first namespace defined by EHLO_Dialect derived from Dialect is called "ehlo", Figure 5 Here, ehlo represents the name of the first namespace corresponding to the first-level transformation and optimization 331; the second namespace defined by EIR_Dialect, which is derived from Dialect, is called "eir," and eir represents the name of the second namespace corresponding to the second-level transformation and optimization 332. For example, the specific process can be referred to the description of step S20.

[0141] Afterwards, executable instructions 350 corresponding to each processing unit are obtained according to the second intermediate representation 340. For example, the specific process can refer to the description of step S30.

[0142] For example, Figure 5 The compilation method shown includes two levels of conversion and optimization. Of course, in practice, more levels of conversion and optimization can be included. These levels of conversion and optimization are connected after the second level of conversion and optimization 332, and corresponding hardware platform feature information is added as needed.

[0143] For example, the following uses the activation function RELU and heterogeneous device in the operation information (computation graph) as Figure 1 Take the structure shown in the figure as an example, and explain it in detail. Figure 5 The activation function operator described in the two-level transformation and optimization in

[15] .

[0144] The activation function is defined as shown in the following formula (1). As the activation layer of a neural network, the activation function defines the nonlinear output of the neuron after linear transformation. As can be seen from formula (1), for the input data x, the result of the activation function is the larger value between x and 0.

[0145]

[0146] Figure 6A This is a schematic diagram of a first intermediate representation provided by an embodiment of the present disclosure. It should be noted that, Figure 6A Only the part related to the activation function is shown, and the first intermediate representation may also include more content.

[0147] Figure 6A The block 400 in FIG. 4 is a computation in the first intermediate representation, which is equivalent to the concept of a function in the C language.

[0148] like Figure 6A As shown, "ENTRY" 410 is the entry point of the function. The calculation starting with "ENTRY" 410 is similar to the main function in the C language. "ROOT" 460 is the return flag of the function, which is equivalent to return in the C language.

[0149] exist Figure 6A In the command, each line is an instruction that makes up the calculation, and the parameters starting with the percentage sign % represent temporary variables that store the results of the operator calculation.

[0150] For example, in line 420, a tensor Arg_0.1 with a dimension of 1*1*4*4 and a data type of 32-bit floating point is defined. Arg_0.1 is the input data of the activation function.

[0151] In line 430, the constant operator is used to define a constant constant.2 with a value of 0. The constant type is a 32-bit floating point type.

[0152] In line 440, the broadcast operator broadcast is used to construct a tensor broadcast.3 with a dimension of 1*1*4*4 and a data type of 32-bit floating point. The value of each element in the tensor is 0.

[0153] In line 450, the maximum operator is used to compare the input data Arg_0.1 with broadcast.3 for maximum value, and the maximum value between the two is assigned to the parameter maximum.4.

[0154] In line 460, the parameter maximum.4 is type-converted and assigned to the parameter tuple.5 and returned.

[0155] Thus, in the first intermediate representation, calculation 400 completes the calculation function of the activation function RELU, and uses the operator constant, the operator broadcast, and the operator maximum.

[0156] Figure 6B A schematic diagram describing the first namespace provided in one embodiment of the present disclosure.

[0157] For example, the namespace defined by EHLO_Dialect derived from Dialect is called "ehlo", thereby determining the first namespace ehlo.

[0158] like Figure 6BAs shown, part 510 defines a derived class EHLO_Dialect and a first namespace "mlir::ehlo", and the derived class EHLO_Dialect is derived from the Dialect class provided by MLIR.

[0159] Section 520 defines the data types used in the first namespace. For example, EHLO_SInt represents a signed integer, EHLO_UInt represents an unsigned integer, EHLO_Int represents any integer type, and EHLO_Tensor represents a tensor type that inherits from different data types.

[0160] Section 530 defines the class EHLO_Op, which is derived from the base class Op provided by MLIR. The derived class EHLO_Op is specifically described by specifying the Op name "mnemonic" and the constraints "traits" of Op.

[0161] Of course, it needs to be explained that Figure 6B A partial description related to the first namespace is shown. More content may be defined in the first namespace, and this disclosure does not limit this.

[0162] Figure 6C A schematic diagram illustrating a reference operator provided in an embodiment of the present disclosure. For example, the reference operator is represented as EHLO_ReluOp, which is an operator defined in the first namespace ehlo for performing activation function calculations.

[0163] For example, the neural network processing unit directly supports activation function operators.

[0164] Specifically, EHLO_ReluOp( Figure 6C 610) is derived from EHLO_Op, and the reference operator name is specified as "Relu" in the specialization list 620. The constraint "[NoSideEffect]" specifies that the reference operator is constrained to have no side effects. The annotation of the reference operator EHLO_ReluOp is described by "summary" 630, and a more detailed description of the reference operator EHLO_ReluOp is given by "decription" 640. The parameters and results required for the reference operator EHLO_ReluOp operation are described by "arguments" 650 and "results" 660. "ins" and "outs" in "arguments" 650 and "results" 660 represent input and output, respectively. The types of the input and output parameters are both tensor types EHLO_Tensor defined in the first namespace.

[0165] After defining the reference operator EHLO_ReluOp in the first namespace ehlo, you can Figure 6A The three operators in the calculation 400 shown are converted into the expression form of the reference operator EHLO_ReluOp.

[0166] Figure 6D A schematic diagram of an intermediate representation of operator fusion provided in one embodiment of the present disclosure.

[0167] It should be noted that Figure 6D Only the parts related to the activation function are shown, and the operator fusion intermediate representation can also include more content.

[0168] like Figure 6D As shown, apply the above Figure 6B and Figure 6C Part of the definition will Figure 6A The calculation 400 in is converted into an expression 700 in the operator fusion intermediate representation. The basic operators in the first intermediate representation (operator constant, operator broadcast, and operator maximum) are converted into activation function operators directly supported by the neural network processing unit.

[0169] In this description, the reference operator is expressed as "ehlo.Relu" in line 710. The input and output parameters required by the reference operator are both tensors with dimensions of 1x1x4x4 and element types of 32-bit floating-point numbers.

[0170] Line 720 is used to complete the data type conversion, assign the activation function calculation result to parameter 1, and return it in line 730.

[0171] Thus, through Figures 6A-6D The first level of conversion and optimization is completed by converting the three atomic-granularity operators (constant, broadcast, maximum) in the first intermediate representation into the expression of the custom reference operator in the first namespace.

[0172] For example, after traversing all operation operators in the first intermediate representation and converting them according to the above process, an operator fusion intermediate representation is obtained.

[0173] Figure 7A A schematic diagram describing the second namespace provided in one embodiment of the present disclosure.

[0174] For example, the namespace defined by EIR_Dialect derived from Dialect is called "eir", and the second namespace is determined to be eir.

[0175] like Figure 7AAs shown, part 810 defines the derived class EIR_Dialect, which is derived from the Dialect class provided by MLIR.

[0176] Section 820 defines the tensor type Tensor and defines an important attribute of Tensor: StorageLevel. StorageLevel describes the storage location of Tensor in heterogeneous devices and is divided into four types:

[0177] 0:ON_THE_FLY – tensors are directly calculated and transferred between modules within the processing unit and are not temporarily saved in any storage area;

[0178] 1: GLOBAL – internal storage area of ​​the neural network processing unit, digital processing unit or central processing unit;

[0179] 2: LOCAL—the first memory 140 in the heterogeneous device 100;

[0180] 3: SYSTEM—the second memory 150 in the heterogeneous device 100 .

[0181] The tensor type Tensor provided in the second namespace eir describes the storage location of each tensor, making it easier to add necessary memory access instructions when the compiler generates a hardware executable program.

[0182] Section 830 defines a tensor type AnyLevelTensor that can be stored in any memory location.

[0183] Figure 7B A schematic diagram describing the second namespace provided in one embodiment of the present disclosure.

[0184] For example, Figure 7B As shown, part 910 defines the class BaseOp, which is the base class of all operators in the second namespace eir and defines the common properties of Op.

[0185] For example, since the neural network processing unit 120 in the heterogeneous device 100 used in this embodiment has a data precision conversion module that allows data of one precision type to be converted to another precision type, the data precision conversion module ENPU_PrecisionCVTAttrs is defined in the second namespace eir. Figure 7B As shown in part 920, the data precision conversion module ENPU_PrecisionCVTAttrs includes three attributes: data truncate, scale, and offset.

[0186] For example, section 930 describes the activation function operator ActivationOp directly supported by the neural network processing unit 120. For example, when the neural network processing unit 120 executes the activation function operator, it is necessary to specify the precision of the input tensor and output tensor of the activation function operator. Therefore, the activation function operator ActivationOp contains two attributes of type ENPU_PrecisionCVTAttrs, "in_cvt" and "out_cvt".

[0187] For example, section 940 describes the reference operator EHLO_ReluOp, which corresponds to the expression ReluOp in the second namespace eir. ReluOp derives from ActivationOp and inherits all of ActivationOp's definitions. Therefore, ReluOp receives a tensor of type AnyLevelTensor and returns a tensor of type AnyLevelTensor. Furthermore, ReluOp includes two input parameters, "in_cvt" and "out_cvt," which are ENPU_PrecisionCVTAttrs attributes.

[0188] For example, Figure 7C A schematic diagram of a second intermediate representation provided by an embodiment of the present disclosure. For example, in the second intermediate representation, the reference operator EHLO_ReluOp is converted to a Relu operator, and the expression of the Relu operator is as follows: Figure 7C For example, the ReLU operator is also the second intermediate operator.

[0189] like Figure 7C As shown, apply the above Figure 7A and Figure 7B Part of the definition will Figure 6D The operator fusion in the intermediate representation converts the expression 700 into the expression form 1000 in the second namespace eir.

[0190] For example, "arg0" represents the input parameter of the Relu operator. Part 1010 indicates the type of the Relu operator input tensor. For example, the input tensor type has a dimension of 1x1x4x4 and an element type of 32-bit floating point numbers, and the GLOBAL in the input tensor type indicates that the Relu operator's operation data needs to be read from the internal storage area of ​​the neural network processing unit. Part 1020 indicates the type of the Relu operator output tensor. For example, the output tensor type has a dimension of 1x1x4x4 and an element type of 32-bit floating point numbers, and the LOCAL in the output tensor type indicates that the Relu operator's operation result will be retained in the first memory 140.

[0191] In this description, the expression form of the Relu operator is “eir.Relu” in line 1030, “arg0” represents the input parameter of the Relu operator, and “%0” represents the output parameter of the Relu operator.

[0192] Part 1040 describes the data accuracy of the input tensor and output tensor during the ReLU operator operation. The data accuracy is determined by the three parameters offset, scale, and truncate.

[0193] Part 1050 describes the return of the ReLU operator's operation result "%0".

[0194] Similarly, through Figures 7A-7C The second level of conversion and optimization is completed by converting the reference operators in the operator fusion intermediate representation into the corresponding expression in the second namespace eir. After traversing all reference operators in the operator fusion intermediate representation and converting them in combination with the above process, the second intermediate representation is obtained.

[0195] according to Figure 5-7C From the description of the above part, we can see that for the activation function in the computational graph, in this implementation scheme, first in the first namespace ehlo, according to the hardware characteristics of the neural network processing unit, the three basic operators that implement the activation function in the first intermediate representation are combined into a reference operator EHLO_ReluOp directly supported by the neural network processing unit; then, in the second namespace eir, combined with the hardware characteristics of the neural network processing unit such as the data storage location and precision conversion, the reference operator EHLO_ReluOp is converted into a Relu operator to give the operator a more specific expression.

[0196] The above discussion uses activation functions as an example to illustrate the process of gradually converting activation functions in a computation graph from the first intermediate representation to the operator fusion intermediate representation and the second intermediate representation. It can be understood that similar conversion, optimization, and definition methods can be used for all processing units, such as digital processing units and central processing units, to convert the first intermediate representation.

[0197] Corresponding to the above-mentioned compilation method, at least one embodiment of the present disclosure further provides a compilation device. Figure 8 A schematic block diagram of a compilation device provided in at least one embodiment of the present disclosure.

[0198] For example, Figure 8 As shown, the compiling device 800 includes: an acquiring unit 801 , a converting unit 802 and a generating unit 803 .

[0199] For example, the compiling device 800 is configured to compile the first intermediate representation based on the graph into executable instructions corresponding to each processing unit.

[0200] For example, the compilation device 800 is applicable to a heterogeneous device, for example, a heterogeneous device including multiple different types of processing units. The specific definition of a heterogeneous device can be referred to the aforementioned compilation method, which will not be repeated here.

[0201] For example, the acquiring unit 801 is configured to acquire a first intermediate representation corresponding to the object to be compiled, wherein the first intermediate representation is a graph-based intermediate representation.

[0202] For example, the conversion unit 802 is configured to perform multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation. For example, the multi-level conversion and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to various processing units. For example, the multi-level conversion and optimization also includes common subexpression elimination, dead code elimination, inferred tensor shape, and other conventional optimizations and conversions.

[0203] For example, the generating unit 803 is configured to obtain executable instructions corresponding to each processing unit according to the second intermediate representation. For example, the executable instructions are used to run on a heterogeneous device.

[0204] For example, the acquisition unit 801, the conversion unit 802, and the generation unit 803 include codes and programs stored in the memory; the processor can execute the codes and programs to implement some or all of the functions of the acquisition unit 801, the conversion unit 802, and the generation unit 803 described above. For example, the acquisition unit 801, the conversion unit 802, and the generation unit 803 can be dedicated hardware devices used to implement some or all of the functions of the acquisition unit 801, the conversion unit 802, and the generation unit 803 described above. For example, the acquisition unit 801, the conversion unit 802, and the generation unit 803 can be a circuit board or a combination of multiple circuit boards used to implement the functions described above. In an embodiment of the present application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-temporary memories connected to the processor; and (3) firmware stored in the memory that can be executed by the processor.

[0205] It should be noted that the acquisition unit 801 is used to implement Figure 3 In step S10 shown, the conversion unit 802 is used to implement Figure 3 In step S20 shown, the generating unit 803 is used to implement Figure 3 The step S30 shown in FIG. Therefore, the specific description of the acquisition unit 801 can refer to the embodiment of the above-mentioned compilation method. Figure 3 For the description of step S10 shown in FIG. 1 , the specific description of the conversion unit 802 can be referred to in the embodiment of the above-mentioned compilation method. Figure 3For the description of step S20 shown in FIG. 8 , the specific description of the generating unit 803 can be referred to in the embodiment of the above-mentioned compilation method. Figure 3 In addition, the compiling device can achieve similar technical effects as the aforementioned compiling method, which will not be described in detail here.

[0206] At least one embodiment of the present disclosure further provides an electronic device, Figure 9 A schematic block diagram of an electronic device provided in accordance with at least one embodiment of the present disclosure.

[0207] For example, Figure 9 As shown, the electronic device includes a processor 901, a communication interface 902, a memory 903, and a communication bus 904. The processor 901, the communication interface 902, and the memory 903 communicate with each other via the communication bus 904. The processor 901, the communication interface 902, the memory 903, and other components can also communicate with each other via a network connection. The present disclosure does not limit the type and function of the network.

[0208] For example, the memory 903 is configured to non-transiently store computer-executable instructions. When the processor 901 is configured to execute the computer-executable instructions, the computer-executable instructions are executed by the processor 901 to implement the compilation method according to any of the above-described embodiments. The specific implementation and related explanations of each step of the compilation method can be found in the above-described embodiments of the compilation method and are not further described here.

[0209] For example, the processor 901 executes the program stored in the memory 903 to implement the compilation method, which is the same as the implementation method mentioned in the embodiment of the compilation method, and will not be repeated here.

[0210] For example, the communication bus 904 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industrial Standard Architecture (EISA) bus. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0211] For example, the communication interface 902 is used to implement communication between the electronic device and other devices.

[0212] For example, the processor 901 and the memory 903 may be provided on a server side (or in the cloud).

[0213] For example, the processor 901 can control other components in the electronic device to perform desired functions. The processor 901 can be a central processing unit (CPU), a network processor (NP), etc., and can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The central processing unit (CPU) can be an X86 or ARM architecture, etc.

[0214] For example, the memory 903 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer-executable instructions may be stored on the computer-readable storage medium, and the processor 901 may execute the computer-executable instructions to implement various functions of the electronic device. Various applications and various data may also be stored in the storage medium.

[0215] For example, for a detailed description of the process of the electronic device executing the compilation, reference may be made to the relevant description in the embodiment of the compilation method, and repeated parts will not be repeated.

[0216] Figure 10 A schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. Figure 10 As shown, one or more computer executable instructions 1002 may be non-transitory stored on a storage medium 1001. For example, when the computer executable instructions 1002 are executed by a processor, one or more steps in the compilation method described above may be performed.

[0217] For example, the storage medium 1001 may be applied to the above-mentioned electronic device and / or the compiling device 800. For example, the storage medium 1001 may include the memory 903 in the electronic device.

[0218] For example, the description of the storage medium 1001 may refer to the description of the memory in the embodiment of the electronic device, and the repeated parts will be omitted.

[0219] Regarding this disclosure, the following points need to be explained:

[0220] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.

[0221] (2) For the sake of clarity, the thickness and size of layers or structures in the drawings used to describe the embodiments of the present invention are exaggerated. It will be understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element can be "directly on" or "under" the other element, or intervening elements may be present.

[0222] (3) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.

[0223] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.

Claims

1. A compilation method, applicable to heterogeneous devices, wherein: The heterogeneous device includes multiple different types of processing units, The compilation method comprises: Obtaining a first intermediate representation corresponding to the object to be compiled, wherein the first intermediate representation is a graph-based intermediate representation; Performing multi-level transformation and optimization on the first intermediate representation to obtain a second intermediate representation; Obtaining executable instructions corresponding to each processing unit according to the second intermediate representation; The multi-level conversion and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to the various processing units; The first intermediate representation includes multiple operators, and the multiple operators are atomic-granularity operators. Performing multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation includes: Obtain at least one reference operator, wherein each reference operator is a basic operator directly supported by at least one processing unit, and the reference operator includes an operator not supported by the first intermediate representation; Converting the plurality of operation operators according to the at least one reference operator to obtain an operator fusion intermediate representation; According to the hardware platform characteristics of each processing unit, the operator fusion intermediate representation is converted and optimized at least one level, and the hardware platform characteristic information related to the multiple processing units is added level by level to obtain the second intermediate representation. The hardware platform characteristic information includes the hardware operation information or hardware optimization information required for the multiple processing units to run each operator.

2. The compiling method according to claim 1, wherein: Converting the multiple operation operators according to the at least one reference operator to obtain an operator fusion intermediate representation includes: Determining at least one operator that performs a function of a target reference operator from the plurality of operators, wherein the target reference operator is any one of the at least one reference operator; The at least one operation operator is converted into an expression form of the target reference operator to obtain the operator fusion intermediate representation.

3. The compiling method according to claim 2, wherein: Converting the at least one operation operator into an expression form of the target reference operator to obtain the operator fusion intermediate representation includes: Determine the first namespace; The at least one operation operator is converted into an expression form of the target reference operator in the first namespace to obtain the operator fusion intermediate representation.

4. The compiling method according to claim 3, wherein: Determine the first namespace, including: Obtaining the first namespace based on a multi-level intermediate representation architecture; The at least one reference operator is defined in the first namespace.

5. The compiling method according to claim 4, wherein: Defining the at least one reference operator in the first namespace includes: The description of each reference operator, the operation parameters and parameter types of each reference operator, and the operation results and result types of each reference operator are defined in the first namespace.

6. The compiling method according to claim 1, wherein: According to the hardware platform characteristics of each processing unit, the operator fusion intermediate representation is converted and optimized at least one level, and hardware platform characteristic information related to the multiple processing units is added level by level to obtain the second intermediate representation, including: Through the at least one level of conversion and optimization, a description directly related to the hardware platform characteristics of each processing unit is added to the operator fusion intermediate representation to obtain the second intermediate representation, wherein the second intermediate representation contains the hardware platform characteristic information of each processing unit.

7. The compiling method according to claim 6, wherein: By means of the at least one level of conversion and optimization, a description directly related to the hardware platform characteristics of each processing unit is added to the operator fusion intermediate representation to obtain the second intermediate representation, including: Determining at least one second namespace corresponding one-to-one to the at least one level of transformation and optimization; Converting the plurality of first intermediate operators included in the operator fusion intermediate representation into expressions in corresponding second namespaces step by step, wherein the expression in the corresponding second namespace of each first intermediate operator includes hardware platform characteristic information of a processing unit that executes each first intermediate operator; The intermediate representation obtained by the final level of transformation and optimization is used as the second intermediate representation.

8. The compiling method according to claim 7, wherein: Determining at least one second namespace corresponding to the at least one level of transformation and optimization includes: Obtaining the at least one second namespace based on a multi-level intermediate representation architecture; In each second namespace, expressions corresponding to the plurality of first intermediate operators are defined, wherein the expression corresponding to each first intermediate operator in each second namespace is added with hardware platform characteristic information of the processing unit executing the first intermediate operator.

9. The compiling method according to claim 8, wherein: In each second namespace, the expressions corresponding to the multiple first intermediate operators are defined, including: For each first intermediate operator, determining a processing unit that executes the first intermediate operator; According to the processing unit, adding an execution unit identifier to the first intermediate operator; Determining hardware platform characteristic information of the processing unit, wherein the hardware platform characteristic information includes hardware operation information or hardware optimization information, the hardware operation information includes one or more of data shape, data precision, and data storage location corresponding to the processing unit, and the hardware optimization information includes one or more combinations of hardware storage analysis, hardware power consumption minimization, data blocking, and hardware storage quantization; According to the hardware platform characteristic information of the processing unit, an expression form corresponding to the first intermediate operator is defined in each second namespace, wherein different hardware platform characteristic information is defined in different second namespaces.

10. The compiling method according to any one of claims 1 to 9, wherein: The second intermediate representation includes a plurality of second intermediate operators, Obtaining executable instructions corresponding to each processing unit according to the second intermediate representation includes: Determining, based on the second intermediate representation, a processing unit corresponding to each second intermediate operator, wherein each second intermediate operator is executed on the corresponding processing unit, and the second intermediate representation includes an execution unit identifier corresponding to each second intermediate operator; According to the instruction set type and rules of each processing unit, hardware platform characteristic information of at least one second intermediate operator executed on each processing unit is extracted from the second intermediate representation to generate executable instructions corresponding to each processing unit.

11. The compiling method according to any one of claims 1 to 9, wherein: Obtaining executable instructions corresponding to each processing unit according to the second intermediate representation includes: Converting the second intermediate representation into an underlying virtual intermediate representation; Generate executable instructions corresponding to each processing unit according to the underlying virtual intermediate representation.

12. The compiling method according to any one of claims 1 to 9, wherein: Obtaining the first intermediate representation corresponding to the object to be compiled includes: Obtaining operation information of the object to be compiled, wherein the operation information is in the form of a calculation graph of the object to be compiled; The object to be compiled is converted into the first intermediate representation according to the operation information.

13. The compiling method according to claim 12, wherein: Converting the object to be compiled into the first intermediate representation according to the operation information includes: Compiling the operation information through an accelerated linear algebra compiler to obtain a high-order optimized intermediate representation; The high-order optimized intermediate representation is used as the first intermediate representation.

14. The compiling method according to any one of claims 1 to 9, wherein: The object to be compiled is used to perform neural network calculations.

15. The compiling method according to any one of claims 1 to 9, wherein: The multiple different types of processing units include at least two of a central processing unit, a graphics processing unit, a neural network processing unit, and a digital processing unit.

16. A compilation device, applicable to heterogeneous devices, wherein: The heterogeneous device includes multiple different types of processing units, The compiling device comprises: an acquiring unit configured to acquire a first intermediate representation corresponding to the object to be compiled, wherein the first intermediate representation is a graph-based intermediate representation; a conversion unit configured to perform multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation; a generating unit configured to obtain executable instructions corresponding to each processing unit according to the second intermediate representation, The multi-level conversion and optimization includes performing operator fusion on the first intermediate representation and adding hardware platform characteristic information related to the various processing units; The first intermediate representation includes multiple operators, and the multiple operators are atomic-granularity operators. Performing multi-level conversion and optimization on the first intermediate representation to obtain a second intermediate representation includes: Obtain at least one reference operator, wherein each reference operator is a basic operator directly supported by at least one processing unit, and the reference operator includes an operator not supported by the first intermediate representation; Converting the plurality of operation operators according to the at least one reference operator to obtain an operator fusion intermediate representation; According to the hardware platform characteristics of each processing unit, the operator fusion intermediate representation is converted and optimized at least one level, and the hardware platform characteristic information related to the multiple processing units is added level by level to obtain the second intermediate representation. The hardware platform characteristic information includes the hardware operation information or hardware optimization information required for the multiple processing units to run each operator.

17. An electronic device comprising: a memory that non-transitorily stores computer-executable instructions; a processor configured to execute the computer-executable instructions, The computer executable instructions, when executed by the processor, implement the compilation method according to any one of claims 1 to 15.

18. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer-executable instructions are executed by a processor, the compilation method according to any one of claims 1 to 15 is implemented.

Citation Information

Patent Citations

  • Neural network compiling method for storage and calculation integrated platform

    CN112465108A

  • Business processing method and device and equipment

    CN113961267A