Operator plug-in type compiling method, device and equipment and storage medium

By using an operator plug-in compilation method, the problem of frequent modifications to the compiler core code required for NPU operator compilation in existing technologies is solved. This achieves efficient operator compilation and hardware adaptation, supports multi-team collaborative development, and improves compilation efficiency and deployment flexibility.

CN120909599APending Publication Date: 2025-11-07CCORE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511102676.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, each new NPU operator requires modification of the compiler core code, resulting in long development and debugging cycles and high risks. The tightly coupled architecture means that introducing new operators or optimizing existing operators requires intrusive modifications to the compiler core.

Method used

The method employs an operator plug-in compilation approach. By obtaining the intermediate representation of the neural network model and the standardized operator description, optimization operations are performed using optimization plug-ins. Pre-created operator compilation plug-ins are loaded and registered to the compilation framework, corresponding operator instantiation objects are created, and compilation is performed to convert them into hardware instructions.

Benefits of technology

It avoids frequent modifications to the core code of the compilation framework, improves operator compilation efficiency, supports collaborative development by multiple teams, and eliminates the need for framework developers to get involved in the underlying details. Operator compilation plugin developers can focus on operator compilation and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909599A_ABST
    Figure CN120909599A_ABST
Patent Text Reader

Abstract

The invention discloses an operator plug-in type compiling method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining an intermediate representation of a neural network model and a standardized operator descriptor corresponding to each operator in the intermediate representation, and performing optimization operation on the intermediate representation by using an optimization plug-in based on the standardized operator descriptor to obtain an optimized intermediate representation; loading a pre-created operator compiling plug-in, registering the operator compiling plug-in to a compiling framework, and creating a corresponding operator instantiation object for each operator in the optimized intermediate representation by using the operator compiling plug-in; and executing compiling through the operator instantiation object to convert the operator in the optimized intermediate representation into a hardware instruction. The frequent modification of core codes of a compiling framework can be avoided, and the operator compiling efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to an operator plug-in type compiling method and device, equipment and a storage medium. BACKGROUND

[0002] With the continuous development of artificial intelligence, a neural network processor (NPU, Neural Network Processing Unit) becomes a core engine supporting intelligent computing. In order to pursue extreme performance and energy efficiency, the NPU architecture design presents a trend of high customization and rapid iteration, and various new operators emerge in an endless stream. How to realize efficient compilation of the operators is a very key step. NPU operator compilation refers to the process of converting operators in a neural network model into an executable instruction set of an NPU chip. This process involves converting operators (such as convolution, matrix operation, etc.) in a high-level programming model into underlying code supported by NPU hardware, and optimizing execution efficiency. In the prior art, each newly added NPU operator needs to modify the core code of the compiler, that is, the tightly coupled architecture makes it necessary to modify the core of the compiler in an invasive manner when introducing new operators or optimizing existing operators, resulting in a long development and debugging cycle and high risk. SUMMARY

[0003] Therefore, the purpose of the present application is to provide an operator plug-in type compiling method, device, equipment and storage medium, which can avoid frequent modification of the core code of the compiling framework and improve the efficiency of operator compilation. The specific scheme is as follows:

[0004] In a first aspect, the present application discloses an operator plug-in type compiling method, comprising:

[0005] obtaining an intermediate representation of a neural network model and a standardized operator description body corresponding to each operator in the intermediate representation, performing an optimization operation on the intermediate representation based on the standardized operator description body by using an optimization plug-in, and obtaining an optimized intermediate representation;

[0006] loading a pre-created operator compiling plug-in and registering the operator compiling plug-in to a compiling framework, and creating a corresponding operator instantiation object for each operator in the optimized intermediate representation by using the operator compiling plug-in;

[0007] performing compilation through the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions.

[0008] Optionally, the obtaining of the intermediate representation of the neural network model and the standardized operator description body corresponding to each operator in the intermediate representation comprises:

[0009] parsing a computation graph structure of the neural network model, traversing all operator nodes in the model and extracting operator basic information;

[0010] The standardization operator description body is obtained by correcting the semantic difference between operators, and the standardization operator description body includes operator type, operator name, topology connection information, and attribute information.

[0011] Optionally, the optimization plug-in is generated by decomposing a layer optimization process into independent optimization steps, and the optimization plug-in is dynamically loaded.

[0012] Optionally, the pre-created operator compilation plug-in is loaded and registered to the compilation framework, including:

[0013] The target operator type corresponding to the currently loaded operator compilation plug-in is determined.

[0014] It is checked whether the target operator type corresponds to an operator compilation plug-in in the loaded operator compilation plug-in.

[0015] If not, the currently loaded operator compilation plug-in is registered to the compilation framework.

[0016] Optionally, the operator instantiation object corresponding to each operator in the optimized intermediate representation is created by using the operator compilation plug-in, including:

[0017] The compilation framework traverses the operators in the topology structure of the intermediate representation, and compares the traversed operators with the operator compilation plug-ins registered in the compilation framework.

[0018] If there is an operator compilation plug-in matching the operator, the operator instantiation object corresponding to the operator is created by using the matched operator compilation plug-in.

[0019] When all operators are matched to the corresponding operator compilation plug-ins, it is checked whether there is an unmatched plug-in in all the operator compilation plug-ins registered in the compilation framework, and if there is, the plug-in is deleted.

[0020] If there is no operator compilation plug-in matching the operator, the operator instantiation object creation process is exited.

[0021] Optionally, the operator instantiation object includes an operator compilation interface.

[0022] The compilation is performed by using the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions, including:

[0023] The compilation is performed by calling the operator compilation interface of the operator instantiation object.

[0024] After all the calling operator compilation interfaces are executed, a deployment file of the neural network model on a to-be-deployed platform is generated based on the hardware instructions, parameter data, and platform verification data generated in the editing stage.

[0025] Optionally, the operator instantiation object comprises a memory application interface, and the memory application interface is configured to apply for a memory resource of the neural network processor.

[0026] Before the operator instantiation object executes the compilation, the method further comprises:

[0027] The memory application request sent by the operator instantiation object through the memory application interface is received.

[0028] Memory resources are allocated to each of the operator instantiation objects according to a global memory planning strategy based on all the memory application requests, and a resource allocation result is returned to the corresponding operator instantiation object through a memory handle; the global memory planning strategy comprises a memory planning based on a use classification and a memory planning based on a life cycle division.

[0029] In a second aspect, the present application discloses an operator plug-in compilation device, comprising:

[0030] An intermediate representation optimization module is configured to obtain an intermediate representation of a neural network model and a standardized operator description body corresponding to each operator in the intermediate representation, perform an optimization operation on the intermediate representation based on the standardized operator description body by using an optimization plug-in, and obtain an optimized intermediate representation.

[0031] An operator instantiation module is configured to load a pre-created operator compilation plug-in, register the operator compilation plug-in to a compilation framework, and create a corresponding operator instantiation object for each operator in the optimized intermediate representation by using the operator compilation plug-in.

[0032] An operator compilation module is configured to perform a compilation by using the operator instantiation object, so as to convert the operators in the optimized intermediate representation into hardware instructions.

[0033] In a third aspect, the present application discloses an electronic device, comprising:

[0034] A memory is configured to save a computer program.

[0035] A processor is configured to execute the computer program, so as to implement the operator plug-in compilation method.

[0036] In a fourth aspect, the present application discloses a computer readable storage medium configured to store a computer program; when the computer program is executed by a processor, the operator plug-in compilation method is implemented.

[0037] In the present application, the intermediate representation of the neural network model and the standardized operator description body corresponding to each operator in the intermediate representation are obtained, the optimization plug-in is used to perform optimization operation on the intermediate representation based on the standardized operator description body, and an optimized intermediate representation is obtained; the operator compilation plug-in created in advance is loaded and registered to the compilation framework, and the operator compilation plug-in is used to create a corresponding operator instantiation object for each operator in the optimized intermediate representation; and the operator instantiation object is used for execution compilation to convert the operators in the optimized intermediate representation into hardware instructions.

[0038] By loading the operator compilation plug-in to register to the compilation framework, operator plug-in type compilation is realized, so that the operator can be developed and optimized for specific hardware, and integrated into the compilation framework in the form of plug-in, without the need to modify the compilation framework every time the operator is modified or added, avoiding frequent modification of the core code of the compilation framework and improving the efficiency of operator compilation; and plug-in type compilation is conducive to multi-team collaborative development, and the framework developer does not need to intervene in the underlying details, while the operator compilation plug-in developer can focus on operator compilation and optimization. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0040] Figure 1 An operator plug-in type compilation method flowchart is provided for the present application;

[0041] Figure 2 A specific operator plug-in type compilation method flowchart is provided for the present application;

[0042] Figure 3 An operator plug-in type compilation device structure schematic diagram is provided for the present application;

[0043] Figure 4 An electronic equipment structure diagram is provided for the present application. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0045] In the prior art, each newly added NPU operator needs to modify the core code of the compiler, that is, the tightly coupled architecture makes it necessary to modify the core of the compiler when introducing new operators or optimizing existing operators, resulting in a long development and debugging period and high risk. To overcome the above technical problems, the present application proposes an operator plug-in type compilation method, which can avoid frequent modification of the core code of the compilation framework and improve the efficiency of operator compilation.

[0046] The embodiment of the present application discloses an operator plug-in type compilation method, as shown in Figure 1 The method can include the following steps:

[0047] Step S11: Obtain the intermediate representation of the neural network model and the standardized operator description body corresponding to each operator in the intermediate representation, and perform optimization operation on the intermediate representation based on the standardized operator description body using the optimization plug-in to obtain the optimized intermediate representation.

[0048] Intermediate representation (IR) is a key abstraction layer connecting algorithm models and hardware execution environments. In the construction process of IR, each computing operator is parsed into a standardized operator description body (Standardized Operator Descriptor); the description body defines at least the following core elements: operator type, tensor descriptor (including shape, data type, memory layout), computing attribute (such as convolution kernel size, step parameter) and data dependency relationship. The standardized metadata representation ensures that the operator semantics of different front-end frameworks (PyTorch / TensorFlow) can be parsed and processed under a unified paradigm.

[0049] Based on this standardized description system, the optimization engine implements multi-level IR conversion through a plug-in architecture. The optimization plug-in executes compile-time optimizations including operator fusion, constant folding, data layout transformation, etc. according to the standardized interface provided by the description body. Each optimization stage strictly maintains the computational semantic equivalence, while specific improvements are made for the target hardware characteristics. After multiple rounds of optimization iterations, the IR not only retains the functional integrity of the original model, but also significantly improves the execution efficiency on the target computing device, laying an optimization foundation for the subsequent code generation stage. This standardized processing flow effectively solves the hardware adaptation challenge in deep learning model deployment, and realizes efficient mapping from algorithm to hardware.

[0050] In a specific implementation, the obtaining of the intermediate representation of the neural network model and the standardized operator description corresponding to each operator in the intermediate representation can include: parsing the computation graph structure of the neural network model, traversing all operator nodes in the model and extracting operator basic information; correcting the semantic difference between operators to obtain the standardized operator description; the standardized operator description includes operator type, operator name, topology connection information, and attribute information. It can be understood that model parsing is the primary stage of the NPU operator plug-in compilation process, and its main task is to convert the neural network model described by the deep learning framework into an intermediate representation that can be processed by the subsequent stage. The present embodiment realizes ONNX (pen Neural Network Exchange, open neural network exchange) full-quantity operator support, parses the model topology structure through deserialization, and automatically processes the semantic difference between operators of different versions to generate a standardized operator description body, which includes operator type, operator name, topology connection information, attributes, and other key information, providing a structured query interface for subsequent plug-in matching and operator compilation.

[0051] For example, the original model is converted into a hardware-independent IR representation through parsing, and the IR completely retains the operator nodes, data flow topology, and calculation semantics of the model. In the IR generation stage, each operator node is further parsed into a standardized operator description body, which not only includes operator type (such as Conv2D, MatMul), shape and data type of input and output tensors, and other basic information, but also includes calculation parameters (such as stride and padding values of convolution) and topology connection relationship of the operator in the computation graph. The standardized description body effectively solves the difference problem in operator implementation of different deep learning frameworks (such as PyTorch and TensorFlow), and provides a unified interface specification for subsequent optimization. Based on the standardized description body, the compilation framework can intelligently select and call various optimization plugins, such as layer optimization plugins that can perform operator fusion (combining Conv2D and ReLU into a single operator), data layout conversion (NCHW (Batch-Channel-Height-Width, data organization method arranged according to batch size, channel number, height, and width) to NHWC (Batch-Height-Width-Channel, data organization method arranged according to sample number, height, width, and channel) optimization operations; and operator-level optimization plugins that generate highly optimized calculation instructions for specific hardware; optimization operations can significantly improve the execution efficiency of the model on the target hardware under the premise of maintaining calculation semantic equivalence.

[0052] In a specific implementation, the optimization plugin is a dynamically loadable optimization plugin generated by decomposing the layer optimization process into independent optimization steps. It can be understood that, as a hardware accelerator specially designed for AI computing, the NPU has a fundamental difference in architecture from the traditional CPU / GPU. Layer optimization mainly refers to optimizing the computation graph of an AI model to improve the execution efficiency and performance of the hardware; the core of layer optimization lies in deep adaptation to the computing characteristics and hardware architecture to maximize the hardware potential. However, a single optimization strategy cannot adapt to diversified hardware. The present application decouples the traditional fixed layer optimization process into dynamically combinable optimization units, and different optimization steps of different hardware architectures are integrated in the form of plugins. The layer optimization plugin only needs to include an optimization interface, which takes an intermediate representation as input. The optimization plugin modifies the IR intermediate representation to achieve optimization purposes such as operator combination, operator splitting, data format conversion processing, etc., wherein the modification content can be an attribute of the IR intermediate representation, or the input / output tensor dimensions thereof, or the model topology structure.

[0053] Step S12: loading a pre-created operator compilation plugin and registering the operator compilation plugin to the compilation framework, and creating a corresponding operator instantiation object for each operator in the optimized intermediate representation using the operator compilation plugin.

[0054] In a specific implementation, the loading of the pre-created operator compilation plugin and the registration of the operator compilation plugin to the compilation framework include: determining the target operator type corresponding to the currently loaded operator compilation plugin; checking whether the operator compilation plugin corresponding to the target operator type is included in the loaded operator compilation plugin; if not, registering the currently loaded operator compilation plugin to the compilation framework.

[0055] Specifically, the compilation framework scans the plugins under the compilation plugin directory, and the operator compilation object instantiation method provided by the operator compilation plugin needs to conform to the pre-agreed naming method. If it conforms, the plugin is dynamically loaded and registered to the compilation framework, so that the pre-created operator compilation plugin can be loaded. The registration takes the operator type as the unique key value, and checks whether the operator compilation plugin corresponding to the target operator type is included in the loaded operator compilation plugin, that is, if the type of operator has been registered, the registration behavior is stopped and the plugin is unloaded. The operator type is an identifier that distinguishes different computing operations, which determines which compilation plugin is called; for example, the operator compilation plugin for convolution operation is only registered once, and multiple operator compilation plugins corresponding to sub-operations are not registered to avoid conflicts. The operator compilation plugin not only provides a method for creating an operator compilation instance, but also declares the supported operator types.

[0056] The registration process is essentially to abstract hardware-specific compilation capabilities into framework-schedulable service resources, and the framework maintains a global plug-in registry to achieve centralized management and on-demand allocation of compilation resources. The plug-in registration mechanism eliminates the compilation-time dependency between the framework core and the specific hardware implementation, enabling the framework to dynamically bind the compilation logic of different hardware backends at runtime; through standardized plug-in interface contracts, it ensures that specialized compilers for heterogeneous computing devices can access the framework in a unified manner.

[0057] In the deep learning model compilation process, the implementation of the plug-in architecture involves dynamic loading and registration mechanisms for operator compilation plugins. The loading operation of pre-generated operator compilation plugins encapsulates code generation logic for specific hardware architectures in the form of dynamic link libraries. By integrating various operator compilation plugins into a unified compilation framework environment through standardized operator compilation plugin interface specifications, the system can ensure that compilation plugins provided by different hardware vendors can work collaboratively under the same framework.

[0058] After the operator compilation plugin is registered, the compilation framework performs operator instantiation based on the optimized intermediate representation. This stage creates corresponding instantiation objects for each operator node in the intermediate representation through the registered operator compilation methods. These instantiation objects serve as specific execution units for the operator compilation process, encapsulating target hardware-related compilation context information, including instruction generation strategies, memory allocation schemes, and hardware-specific optimization parameters. Creating operator instantiation objects enables fine-grained compilation context isolation, with each instantiation object independently encapsulating the compilation state of an operator. The isolation mechanism supports parallel compilation of multiple operators, avoiding global state contention. Operator instantiation objects are generated through the operator compilation methods provided by operator compilation plugins. This implementation achieves modularity and configurability in the compilation process, allowing the framework to maintain stability in core logic while flexibly adapting to evolving hardware architectures. Through plugin registration and instantiation mechanisms, the framework maximizes the computational potential of hardware while ensuring compilation correctness.

[0059] In specific implementations, the use of the operator compilation plugin to create corresponding operator instantiation objects for each operator in the optimized intermediate representation includes: the compilation framework traverses the operators in the topology structure of the intermediate representation and compares the traversed operators with the registered operator compilation plugins in the compilation framework; if there is an operator compilation plugin that matches an operator, the matched operator compilation plugin is used to create a corresponding operator instantiation object for the operator; after all operators have matched the corresponding operator compilation plugins, it is checked whether there are any unmatched plugins among all the registered operator compilation plugins in the compilation framework, and if there are, the plugin is deleted; if there is no operator compilation plugin that matches the operator, the operator instantiation object creation process is exited.

[0060] When the operator compilation plug-in is loaded, the compilation framework starts the operator compilation object instantiation, the instantiated operator compilation object is held by the corresponding management module of the compilation framework, and is destroyed after the compilation ends, while releasing its associated memory resources, etc. The method of operator compilation object instantiation is that the compilation framework traverses the intermediate representation topology, matches the operator compilation plug-in for each operator in the intermediate representation topology, and the operator compilation plug-in has completed scanning, registration and loading in the foregoing process. The matching rule can be full-word case-insensitive matching of operator type, if the operator compilation plug-in is found, the plug-in is used to instantiate the operator compilation object, and is bound with the intermediate representation; if the matching fails, the operator compilation object instantiation is exited and the subsequent process is not continued. When each operator in the intermediate representation topology is successfully matched, all loaded operator compilation plug-ins are traversed, and it is checked whether the provided operator compilation method is referenced, if not, the plug-in is unloaded to reduce the memory occupation of the compilation tool. Operator instantiation can isolate the compilation state, the same plug-in can be called by multiple operators of the same type, but the input shape and parameters of each operator can be different, and an independent compilation context is needed. At the same time, different operators can generate code in parallel after instantiation, which improves the compilation speed.

[0061] Step S13: performing compilation through the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions.

[0062] In a specific implementation, the operator instantiation object includes a memory application interface, and the memory application interface is configured to apply for memory resources of the neural network processor; before performing the compilation through the operator instantiation object, the method further includes: receiving a memory application request sent by the operator instantiation object through the memory application interface; allocating memory resources for each operator instantiation object according to a global memory planning strategy based on all memory application requests, and returning a resource allocation result to the corresponding operator instantiation object through a memory handle; and the global memory planning strategy includes a memory planning based on a use classification and a memory planning based on a life cycle division.

[0063] The operator instantiation object is used for memory management requirements, memory application (input / output buffer, workspace) of each operator is independent, and the demand can be accurately declared after instantiation, which facilitates unified planning of the framework. The instantiation object serves as an execution unit of the compilation process, accurately manages the occupied memory, registers and other hardware resources, and releases the instance in time after compilation to realize efficient recovery of resources. The decoupling of the compilation logic and the hardware implementation is realized through the operator compilation plug-in, which enables the framework to flexibly support heterogeneous computing devices; each instantiation object independently manages its compilation state and interacts with the framework core through a standard interface, which ensures the controllability of the compilation process and maintains the expansion capability of the system. The adaptation efficiency of the compilation system to new hardware platforms is significantly improved, providing a foundation for cross-platform model deployment.

[0064] In a specific implementation, the operator instantiation object includes an operator compilation interface; the compilation is performed through the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions, including: performing compilation by calling the calling operator compilation interface of the operator instantiation object; after the execution of all the calling operator compilation interfaces is completed, generating a deployment file of the neural network model on a to-be-deployed platform based on the hardware instructions, parameter data and platform verification data generated in the editing stage.

[0065] That is, the operator compilation instantiation object in the embodiment is forced to depend on two clearly defined standard interfaces. First, the memory application interface, which applies for and prepares all memory resources required by the operator instance for its runtime on the target NPU hardware device. These memories include but are not limited to hardware instruction space, memory space of operator input / output tensors, temporary workspace required for internal calculation of the operator, weight data memory space, etc. The memory application interface internally uniformly calls the memory application interface provided by the compilation framework, and each successfully applied memory returns a memory handle. The handle contains the name, size, virtual address and hardware physical address offset of the memory at the time of application. The other is the operator compilation interface, the core function of which is to perform the compilation process of the operator itself, to convert the high-level calculation logic of the operator into low-level instructions optimized for the target hardware and executable. The key information required for the interface calculation, such as input / output tensors, data types, shapes, dimensions, strides, data layouts, attributes or parameters, is provided by the intermediate representation, while the physical address offset required for hardware calculation is provided by the aforementioned memory application interface.

[0066] When all the operator corresponding to the compiled object instantiation is completed (i.e. each operator handles the basic information required to perform its specific computing logic and the above two standard interfaces), the compilation framework formally starts the entire model compilation workflow. The specific process is as follows: first, the first time to traverse the operator compilation instantiation object, call the memory allocation interface of the operator compilation instantiation object, which does not perform underlying memory allocation, but declares memory requirements. When declaring memory requirements, key attributes must be specified to accurately describe memory requirements: such as memory name, memory type (hardware instructions, weight / parameter data, input / output data, temporary workspace), memory size, memory address alignment size, etc.

[0067] After traversing all operator compilation instantiation objects, the compilation framework globally counts and classifies all memory applications collected. The classification is mainly based on memory usage and life cycle, and the specific classification categories include but are not limited to: hardware instructions (store compiled device executable code, life cycle is the same as the entire model), weights / parameters (for example, store the weights data required in the inference process after the model is trained, the life cycle is the same as the entire model), model input / output (store the input data and final output data of the entire model, the model input life cycle is released after being used, and the model output life cycle is the same as the entire model), operator input / output (store the intermediate tensor passed between operators, the life cycle is limited to before and after operator execution), temporary workspace (temporary buffer required for internal computation of operators, the life cycle is limited to the operator execution period); based on the classification, size, address alignment requirements of memory and the memory architecture characteristics of the target hardware, the compilation framework performs global memory planning, and the framework divides a logically or physically continuous memory area for each type of memory; for example, all weights / parameters are placed in a large continuous area, and all operator outputs are placed in another area. In the planned memory area, the memory allocator of the framework performs actual memory allocation. The allocation result is embodied in the memory handle returned by the framework to each memory application request, which contains key information such as virtual address (logical address accessible by operator compilation plug-ins), physical address offset (actual physical address offset of hardware memory), memory size, etc.

[0068] After completing global memory planning and allocation, the second time to traverse the operator compilation instantiation object, for each operator compilation object, the framework calls the operator compilation interface provided by it, which completes the core compilation work. The compilation interface compilation work includes but is not limited to: hardware instruction conversion (converts the operator high-level computing description into the native instruction set of the target hardware platform), data processing (processes data rearrangement based on the specific memory address and layout information obtained), computing optimization (uses the characteristics of the target hardware for deep optimization), etc.

[0069] Finally, when all operator compilations are completed, the compilation framework will generate hardware instruction sets, weight / parameter data, platform verification data (meta-information embedded in the compilation output file, used to verify the compatibility of the generated code with the target platform, record hardware constraints (such as memory alignment requirements, instruction set version) at the time of compilation, etc.) during the compilation phase, and encapsulate them into a single deployment file (bin file) that can be directly distributed and deployed according to the specific format and specification of the target platform (i.e. the hardware environment where the generated code will eventually run, including NPU model, supporting software stack (driver, firmware, runtime library version), memory architecture (such as HBM, DDR configuration), etc.).

[0070] For example Figure 2 As shown in the layer optimization, the compilation framework performs scanning and loading of the operator compilation plug-in, and generates an operator instantiation object using the operator compilation plug-in. The operator instantiation object initiates memory application, and the compilation framework uniformly allocates resources for memory management according to all memory applications, and feeds back the memory resource allocation result to the corresponding operator instantiation object. Finally, the operator compilation interface calling object is called for compilation to generate a target file (i.e. a deployment file).

[0071] The NPU operator plug-in compilation method proposed in the present application aims to fundamentally solve the adaptation problem of the rapid evolution of the compiler and the operator. Through the operator compilation plug-in, multiple vendors' NPUs are supported without modifying the compilation framework code. New operators only need to develop plug-ins without recompiling the framework. Through standardized operator description, differences between front-end operators are eliminated. By creating an independent operator instantiation object for each operator, multi-operator parallel compilation is supported, and NPU memory / register resources are accurately managed. Dynamic expansion and instant integration of operators are realized, improving compilation efficiency and deployment flexibility, and solving the NPU hardware adaptation problem. The plug-in compilation loads the operator compilation module on demand, reducing the memory occupation and startup time of the compilation tool. Moreover, it solves the problem of unclear responsibility boundaries between operator developers and compiler developers, and the difficulty of efficient accumulation and reuse of operator optimization knowledge. It improves the efficiency of NPU innovation landing and application deployment.

[0072] As can be seen, in the embodiment, the intermediate representation of the neural network model and the standardized operator description corresponding to each operator in the intermediate representation are obtained, the intermediate representation is optimized based on the standardized operator description by using an optimization plug-in to obtain an optimized intermediate representation; an operator compilation plug-in created in advance is loaded and registered to a compilation framework, and the operator compilation plug-in is used to create a corresponding operator instantiation object for each operator in the optimized intermediate representation; and the operator instantiation object is used to perform compilation to convert the operators in the optimized intermediate representation into hardware instructions. As can be seen, by loading the operator compilation plug-in and registering it to the compilation framework, operator plug-in type compilation is realized, so that operators can be developed and optimized for specific hardware and integrated into the compilation framework in the form of plug-ins, without the need to modify the compilation framework every time an operator is modified or added, so that the core code of the compilation framework is not frequently modified, and the efficiency of operator compilation is improved; and plug-in type compilation is conducive to multi-team collaborative development, without the need for framework developers to intervene in underlying details, and operator compilation plug-in developers can focus on operator compilation and optimization.

[0073] Correspondingly, the embodiment of the application further discloses an operator plug-in type compilation device, which refers to Figure 3 As shown in the figure, the device comprises:

[0074] An intermediate representation optimization module 11 is configured to obtain an intermediate representation of a neural network model and a standardized operator description corresponding to each operator in the intermediate representation, and to optimize the intermediate representation based on the standardized operator description by using an optimization plug-in to obtain an optimized intermediate representation.

[0075] An operator instantiation module 12 is configured to load an operator compilation plug-in created in advance and register the operator compilation plug-in to a compilation framework, and to use the operator compilation plug-in to create a corresponding operator instantiation object for each operator in the optimized intermediate representation.

[0076] An operator compilation module 13 is configured to perform compilation by using the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions.

[0077] As can be seen, in the embodiment, the operator plug-in type compilation is realized by loading the operator compilation plug-in and registering it to the compilation framework, so that operators can be developed and optimized for specific hardware and integrated into the compilation framework in the form of plug-ins, without the need to modify the compilation framework every time an operator is modified or added, so that the core code of the compilation framework is not frequently modified, and the efficiency of operator compilation is improved; and plug-in type compilation is conducive to multi-team collaborative development, without the need for framework developers to intervene in underlying details, and operator compilation plug-in developers can focus on operator compilation and optimization.

[0078] In some specific embodiments, the intermediate representation optimization module 11 can specifically comprise:

[0079] The parsing unit is configured to parse a computation graph structure of the neural network model, traverse all operator nodes in the model, and extract operator basic information.

[0080] The semantic correction unit is configured to correct semantic differences between operators to obtain the standardized operator description body; and the standardized operator description body includes an operator type, an operator name, topology connection information, and attribute information.

[0081] In some specific embodiments, the optimization plug-in is dynamically loadable by decomposing a layer optimization process into independent optimization steps, and generating an optimization plug-in for each independent optimization step.

[0082] In some specific embodiments, the operator instantiation module 12 can specifically include:

[0083] The type determination unit is configured to determine a target operator type corresponding to the currently loaded operator compilation plug-in.

[0084] The query unit is configured to check whether the target operator type corresponds to an operator compilation plug-in in the loaded operator compilation plug-in.

[0085] The registration unit is configured to register the currently loaded operator compilation plug-in to the compilation framework if the target operator type does not correspond to an operator compilation plug-in.

[0086] In some specific embodiments, the operator instantiation module 12 can specifically include:

[0087] The comparison unit is configured to compare the operator in the topology structure of the intermediate representation with the operator compilation plug-in registered in the compilation framework.

[0088] The instantiation object creation unit is configured to create a corresponding operator instantiation object for the operator by using the matched operator compilation plug-in if the operator matches the operator compilation plug-in.

[0089] The plug-in deletion unit is configured to delete the plug-in if the plug-in exists in all operator compilation plug-ins registered in the compilation framework after all operators are matched to the corresponding operator compilation plug-ins.

[0090] The exit unit is configured to exit the operator instantiation object creation process if the operator does not match the operator compilation plug-in.

[0091] In some specific embodiments, the operator instantiation object includes an operator compilation interface.

[0092] The operator compilation module 13 can specifically include:

[0093] The compiling unit is configured to compile by calling the operator instantiation object to execute the calling operator compiling interface;

[0094] The deployment file generation unit is configured to generate a deployment file of the neural network model on a to-be-deployed platform based on the hardware instructions, the parameter data, and the platform verification data generated in the editing stage after all the calling operator compiling interface executions are completed.

[0095] In some embodiments, the operator instantiation object comprises a memory application interface configured to apply for a memory resource of a neural network processor;

[0096] The operator plug-in compiling apparatus can specifically comprise:

[0097] The request receiving unit is configured to receive a memory application request sent by the operator instantiation object through the memory application interface before the operator instantiation object executes the compiling;

[0098] The memory partitioning unit is configured to allocate a memory resource for each operator instantiation object according to a global memory planning strategy based on all the memory application requests, and return a resource allocation result to the corresponding operator instantiation object through a memory handle; the global memory planning strategy comprises a memory planning based on a use classification and a memory planning based on a life cycle division.

[0099] Further, the embodiment of the present application further discloses an electronic device, referring to Figure 4 The content in the figure cannot be considered as any limitation on the use range of the present application.

[0100] Figure 4 A structural schematic diagram of an electronic device 20 is provided in the embodiment of the present application. The electronic device 20 specifically can comprise at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is configured to store a computer program, the computer program is loaded and executed by the processor 21 to implement the related steps in the operator plug-in compiling method disclosed in any of the preceding embodiments.

[0101] In the embodiment, the power supply 23 is configured to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the present application, which is not specifically limited here; the input / output interface 25 is configured to obtain external input data or output data to the outside, and the specific interface type can be selected according to the specific application needs, which is not specifically limited here.

[0102] In addition, the memory 22 as a carrier for storing resources can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon include an operating system 221, a computer program 222, data 223 including a standardized operator description body, etc., and the storage mode can be temporary storage or permanent storage.

[0103] The operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to realize the operation and processing of the processor 21 on the mass data 223 in the memory 22, and can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the operator plug-in type compiling method executed by the electronic device 20 disclosed in any one of the foregoing embodiments, the computer program 222 can further include a computer program capable of completing other specific work.

[0104] Further, the embodiment of the present application further discloses a computer storage medium, which stores computer executable instructions, and the computer executable instructions are loaded and executed by a processor to realize the operator plug-in type compiling method steps disclosed in any one of the foregoing embodiments.

[0105] Further, the embodiment of the present application further discloses a computer program product, which includes a computer program, and the computer program is executed by a processor to realize the operator plug-in type compiling method steps disclosed in any one of the foregoing embodiments.

[0106] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts of each embodiment can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the related parts can be referred to the method part.

[0107] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0108] Finally, it needs to be pointed out that in this document, relational terms such as first and second and the like can only be intended to distinguish one entity or operation from another entity or operation without necessarily requiring or implying any such actual relationship or order between such entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus including the stated element.

[0109] The operator plug-in compilation method, device, equipment and storage medium provided by the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples in this document. The above example description is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range will be changed; in view of the above, the content of the specification should not be understood as a limitation of the present application.

Claims

1. An operator plug-in compilation method, characterized by, The method comprises the following steps: obtaining an intermediate representation of a neural network model and a standardized operator description corresponding to each operator in the intermediate representation, and performing optimization operation on the intermediate representation by using an optimization plug-in based on the standardized operator description to obtain an optimized intermediate representation; loading a pre-created operator compilation plug-in and registering the operator compilation plug-in to a compilation framework, and creating a corresponding operator instantiation object for each operator in the optimized intermediate representation by using the operator compilation plug-in; performing compilation by using the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions.

2. The method of claim 1, wherein, The method of obtaining an intermediate representation of a neural network model and a standardized operator description corresponding to each operator in the intermediate representation comprises the following steps: analyzing the computation graph structure of the neural network model, traversing all operator nodes in the model and extracting operator basic information; correcting the semantic difference between operators to obtain the standardized operator description; the standardized operator description includes operator type, operator name, topology connection information and attribute information.

3. The method of claim 1, wherein, The optimization plug-in is obtained by decomposing the layer optimization process into independent optimization steps, and a dynamically loadable optimization plug-in is generated for each independent optimization step.

4. The method of claim 1, wherein, The method of loading a pre-created operator compilation plug-in and registering the operator compilation plug-in to a compilation framework comprises the following steps: determining the target operator type corresponding to the currently loaded operator compilation plug-in; checking whether the operator compilation plug-in corresponding to the target operator type is included in the loaded operator compilation plug-in; if not, registering the currently loaded operator compilation plug-in to the compilation framework.

5. The method of claim 1, wherein, The method of creating a corresponding operator instantiation object for each operator in the optimized intermediate representation by using the operator compilation plug-in comprises the following steps: the compilation framework traverses the operators in the topology structure of the intermediate representation, and compares the traversed operators with the operator compilation plug-ins registered in the compilation framework; if there is an operator compilation plug-in matching the operator, an operator instantiation object corresponding to the matched operator compilation plug-in is created for the operator; after all operators are matched to the corresponding operator compilation plug-ins, it is checked whether there is an unmatched plug-in in all the operator compilation plug-ins registered in the compilation framework, and if there is, the plug-in is deleted; if there is no operator compilation plug-in matching the operator, the operator instantiation object creation process is exited.

6. The method of claim 1, wherein, The operator instantiation object includes an operator compilation interface. The method of performing compilation by using the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions comprises the following steps: performing compilation by calling the calling operator compilation interface of the operator instantiation object; after the execution of all the calling operator compilation interfaces is completed, a deployment file of the neural network model on a to-be-deployed platform is generated based on the hardware instructions, parameter data and platform verification data generated in the editing stage.

7. The method of claim 1 to 6, wherein, The operator instantiation object includes a memory application interface, and the memory application interface is used to apply for memory resources of a neural network processor. Before performing compilation by using the operator instantiation object, the following steps are further included: receiving a memory application request sent by the operator instantiation object through the memory application interface; The memory resource is allocated to each operator instantiation object based on a global memory planning strategy according to all memory application requests, and the resource allocation result is returned to the corresponding operator instantiation object through a memory handle.

8. An operator plug-in compilation apparatus characterized by comprising: Comprise: The intermediate representation optimization module is configured to obtain an intermediate representation of the neural network model and a standardized operator description body corresponding to each operator in the intermediate representation, perform an optimization operation on the intermediate representation based on the standardized operator description body using an optimization plug-in, and obtain an optimized intermediate representation. The operator instantiation module is configured to load a pre-created operator compilation plug-in and register the operator compilation plug-in to a compilation framework, and create a corresponding operator instantiation object for each operator in the optimized intermediate representation using the operator compilation plug-in. The operator compilation module is configured to perform compilation through the operator instantiation object to convert the operators in the optimized intermediate representation into hardware instructions.

9. An electronic device, comprising: Comprise: A memory for saving a computer program; A processor for executing the computer program to implement the operator plug-in type compilation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A memory for storing a computer program; wherein the computer program is executed by a processor to implement the operator plug-in type compilation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • OpenVX framework system for NPU acceleration

    CN121560823A