Compilation method, device and electronic equipment

Through the combination of deep learning compiler and target compiler, deep learning models are compiled into target models that can run on GPGPU, solving the problem of configuring dependencies for different model formats and achieving efficient model compilation and reasoning.

CN119621073BActive Publication Date: 2025-05-16龙芯中科(合肥)技术有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510162234.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

When performing model deployment inference, different dependency libraries need to be configured for different model formats, resulting in high compilation complexity and low efficiency.

Method used

By calling the deep learning compiler, converting the deep learning model into an intermediate expression file, and compiling the intermediate expression file using the target compiler that matches the GPGPU architecture, building the running environment of the deep learning model, thereby generating a target model that can run on the GPGPU.

Benefits of technology

It simplifies the model compilation process, reduces the compilation complexity, improves the model compilation efficiency, avoids the problem of conversion of different model formats, and does not need to configure the inference engine specifically for the GPGPU platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621073B_ABST
    Figure CN119621073B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a compilation method, device and electronic device, which are applied to a general-purpose graphics processor, and the method includes: calling a deep learning compiler to convert a deep learning model to be compiled into an intermediate expression file of the deep learning compiler; obtaining a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format; using the target compiler to compile the intermediate expression file, and constructing a running environment corresponding to the deep learning model to obtain a compiled target model; importing the target image data to be processed into the target model for processing to obtain a target processing result. The embodiment of the present invention simplifies the model compilation process, reduces the compilation complexity, and improves the model compilation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a compiling method, device and electronic equipment. Background Art

[0002] The way deep learning models are compiled usually involves specific frameworks and tools that provide interfaces for configuring and optimizing models. During the model compilation process, losses, optimizers, and metrics need to be assigned to the model. The type of loss depends on the type of problem being addressed and the goal. The optimizer refers to the algorithm used to train the model. Metrics are metrics used to evaluate model performance, which can be accuracy or other user-defined metrics.

[0003] The main purpose of model compilation is to generate effective code implementations for the models described by the deep learning framework on various hardware platforms. This compilation process converts the model definition into highly optimized code for a specific hardware architecture, thereby achieving efficient reasoning and computing.

[0004] Different training frameworks will produce different model formats after training, which means that the model needs different dependency libraries when deploying reasoning, and the differences between different versions of the same framework will be relatively large. Summary of the invention

[0005] The embodiments of the present invention provide a compilation method, device and electronic device, which can solve the problem of needing to configure different dependency libraries for different model formats when performing model deployment reasoning.

[0006] In order to solve the above problem, an embodiment of the present invention discloses a compilation method, which is applied to a general-purpose graphics processor and includes:

[0007] Calling a deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler; the intermediate expression file includes functions and parameters in a function set of the internal representation model of the deep learning compiler;

[0008] Obtaining a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format;

[0009] Compile the intermediate expression file using the target compiler, and build an operating environment corresponding to the deep learning model to obtain a compiled target model;

[0010] The target image data to be processed is imported into the target model for processing to obtain a target processing result.

[0011] Optionally, calling a deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler includes:

[0012] Calling the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model;

[0013] Set the input data type object and tensor shape of the deep learning model;

[0014] Based on the large binary file, the data type, and the tensor shape, functions and parameters corresponding to the deep learning model are configured in the intermediate expression file of the deep learning compiler.

[0015] Optionally, compiling the intermediate expression file using the target compiler and constructing a running environment corresponding to the deep learning model to obtain a compiled target model includes:

[0016] Configuring the graph structure, libraries and parameters of the deep learning model in a runtime build script based on the intermediate expression file;

[0017] Execute the runtime build script to generate an operating environment for the deep learning model.

[0018] Optionally, before importing the target image data to be processed into the target model for processing, the method further includes:

[0019] The image to be processed is preprocessed to convert the image to be processed into target image data; the target image data is text data in a target format.

[0020] Optionally, the step of importing the target image data to be processed into the target model for processing to obtain a target processing result includes:

[0021] The target image data to be processed is imported into the target model, and the operating environment is called so that the target model can infer the target image data and output a target processing result.

[0022] On the other hand, an embodiment of the present invention discloses a compiling device, which is applied to a general-purpose graphics processor, and includes:

[0023] A model conversion module, used to call a deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler; the intermediate expression file includes functions and parameters in a function set of the internal representation model of the deep learning compiler;

[0024] A compiler acquisition module, used to acquire a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format;

[0025] A model compilation module, used to compile the intermediate expression file using the target compiler, and build an operating environment corresponding to the deep learning model to obtain a compiled target model;

[0026] The model inference module is used to import the target image data to be processed into the target model for processing to obtain the target processing result.

[0027] Optionally, the model conversion module includes:

[0028] A download submodule, used to call the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model;

[0029] A setting submodule, used to set the input data type object and tensor shape of the deep learning model;

[0030] A configuration submodule is used to configure functions and parameters corresponding to the deep learning model in an intermediate expression file of the deep learning compiler based on the large binary file, the data type, and the tensor shape.

[0031] Optionally, the model compilation module includes:

[0032] A script configuration submodule, for configuring the graph structure, library and parameters of the deep learning model in a runtime build script based on the intermediate expression file;

[0033] The runtime generation submodule is used to execute the runtime construction script to generate the operating environment of the deep learning model.

[0034] Optionally, the device further comprises:

[0035] The preprocessing module is used to preprocess the image to be processed and convert the image to be processed into target image data; the target image data is text data in a target format.

[0036] Optionally, the model compilation module includes:

[0037] The compiling submodule is used to import the target image data to be processed into the target model and call the operating environment so that the target model can infer the target image data and output a target processing result.

[0038] On the other hand, an embodiment of the present invention further discloses an electronic device, which includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors to perform the aforementioned compilation method.

[0039] The embodiment of the present invention further discloses a readable storage medium. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the aforementioned compiling method.

[0040] The embodiments of the present invention include the following advantages:

[0041] The embodiment of the present invention provides a compilation method, which can first call a deep learning compiler to convert a deep learning model into an intermediate expression file of the deep learning compiler, and then use a target compiler that matches the architecture of GPGPU to compile the intermediate expression file, and build an operating environment for the deep learning model, so as to compile the deep learning model into a target model that can run on GPGPU, and use the target model to process target image data on GPGPU. The embodiment of the present invention does not need to consider the conversion problem of different model formats, nor does it need to configure an inference engine specifically for the GPGPU platform. The deep learning model is cross-compiled by the deep learning compiler and the target compiler that matches the GPGPU architecture, which simplifies the model compilation process, reduces the compilation complexity, and improves the model compilation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0043] Figure 1 is a flowchart of steps of an embodiment of a compiling method of the present invention;

[0044] Figure 2 is a structural block diagram of an embodiment of a compiling device of the present invention;

[0045] Figure 3 It is a structural block diagram of an electronic device for compiling provided by an example of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Method Embodiment

[0048] Reference Figure 1 , shows a flowchart of a compilation method embodiment of the present invention, the method is applied to a general-purpose graphics processor, and the method may specifically include the following steps:

[0049] Step 101: calling a deep learning compiler to convert a deep learning model to be compiled into an intermediate expression file of the deep learning compiler; the intermediate expression file includes functions and parameters in a function set of an internal representation model of the deep learning compiler;

[0050] Step 102: Obtain a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format;

[0051] Step 103: compile the intermediate expression file using the target compiler, and build an operating environment corresponding to the deep learning model to obtain a compiled target model;

[0052] Step 104: Import the target image data to be processed into the target model for processing to obtain a target processing result.

[0053] The compilation method provided in the embodiment of the present invention can be applied to general-purpose computing on graphics processing units (GPGPU). GPGPU is a graphics processor that uses graphics processing tasks to calculate general computing tasks originally processed by the central processing unit. These general computing tasks often have nothing to do with graphics processing.

[0054] Among them, the deep learning compiler is used to convert the deep learning model into an intermediate expression file. Among them, the deep learning model may include a floating-point model or a fixed-point model. Specifically, the deep learning compiler in the embodiment of the present invention is used to generate functions and parameters in the intermediate expression file according to the functions and parameters of the deep learning model. Exemplarily, the deep learning compiler in the present invention can be a compiler composed of a compilation interface provided by TVM, which can be a python interface provided by TVM, which is used to convert the deep learning model into functions and parameters in TVM IR. Among them, function can be written into relay_function, and parameters can be written into relay_parameters. relay_function corresponds to an end-to-end model, which can be simply understood as a computational graph that supports control flow, recursion, and complex data structures. Alternatively, the deep learning compiler in the present invention can also be a compiler generated according to the torch_glow interface in Glow, which is used to convert the PyTorch model into an intermediate expression file GlowIR. Alternatively, the deep learning compiler in the present invention may also be a compiler generated according to the compilation interface in XLA (Accelerated Linear Algebra), for example, using the compilation interface in XLA to convert the TensorFlow calculation graph of the deep learning model into an intermediate expression file. Alternatively, the deep learning model in the present invention may also be a compiler generated according to the compilation interface of MLIR (Multi-Level Intermediate Representation), for example, the deep learning compiler in the present invention is generated according to torch-mlir in MLIR, and the deep learning model is converted into the intermediate expression file MLIR, etc. It should be noted that the embodiment of the present invention does not specifically limit the specific generation process and composition of the deep learning compiler, as long as the deep learning model can be converted into the intermediate expression file of the deep learning compiler.

[0055] The target compiler in the embodiment of the present invention is used to translate the deep learning model into an equivalent target model in machine language format.

[0056] It should be noted that the formats of deep learning models vary, mainly depending on the deep learning framework and storage method used. Different deep learning frameworks and tools may support different model formats, such as model formats based on TensorFlow, PyTorch, Caffe, OnnxRuntime and other frameworks. In order for different deep learning model formats to run on various GPGPU devices, model compilation processing is required to generate effective code implementations for the models described by the deep learning framework on various hardware platforms so that they can run smoothly on the corresponding GPGPU devices. This compilation process converts the model definition into highly optimized code for a specific hardware architecture, thereby achieving efficient reasoning and computing.

[0057] In an embodiment of the present invention, a target compiler is used to convert a deep learning model into a target model so that the target model can be run on a general-purpose graphics processor for image processing. It can be understood that the language format of the compiled target model matches the architecture of the general-purpose graphics processor.

[0058] Exemplarily, the target compiler in the embodiment of the present invention may be a compiler generated according to the compiler architecture of LLVM (Low Level Virtual Machine), which is used to compile the intermediate expression file of the deep learning model, convert the deep learning model represented by the high-level language into machine code, and build the operating environment corresponding to the deep learning model to obtain the compiled target model. Alternatively, the target compiler in the present invention may also be a compiler generated according to the compilation interface of XLA, which is used to compile the intermediate expression file HLO of XLA into machine code, and build the corresponding operating environment to obtain the target model. Alternatively, the target compiler in the present invention may also be a compiler generated according to the compilation interface in OpenVINO, which is used to compile the intermediate expression file OpenVINO IR corresponding to the TensorFlow model, and build the corresponding operating environment to obtain the compiled target model, and so on. Similarly, the embodiment of the present invention does not specifically limit the specific generation process and composition of the target compiler, as long as the deep learning model can be translated into an equivalent target model in machine language format.

[0059] It should be noted that different training frameworks will produce different model formats after training, which means that different dependency libraries are required when the model is deployed for inference, and the differences between different versions of the same framework (such as TensorFlow) will be relatively large.

[0060] In related technologies, a more general reasoning framework (OnnxRuntime, ORT) is usually used for cross-platform machine learning model reasoning, supporting multiple programming languages ​​and frameworks, operating systems and hardware platforms, including GPGPU. After converting the models of frameworks such as PyTorch, TensorFlow, and scikit-learned into ONNX models, the model reasoning can be performed using the OnnxRuntime reasoning engine based on the GPGPU platform, without the need for the original training framework. This makes the deployment of the model more convenient and universal. In addition, OnnxRuntime can achieve faster reasoning speeds through built-in graph optimization strategies and integrated hardware acceleration libraries. Even on the same hardware platform, OnnxRuntime can run faster than PyTorch and TensorFlow.

[0061] However, the converted ONNX format model may have operator support differences on the GPGPU platform, resulting in the model not being able to run normally; in addition, the corresponding OnnxRuntime framework needs to be adapted for GPGPU.

[0062] In an embodiment of the present invention, the deep learning model can be cross-compiled on the GPGPU platform through the deep learning compiler and the architecture compiler, so that the compiled deep learning model can be run on the GPGPU platform to perform data processing.

[0063] Specifically, first, the deep learning compiler is called to convert the deep learning model to be compiled into an intermediate representation file of the deep learning compiler. The intermediate representation file (IR) is used to express various parameters and inter-layer relationships of the deep learning model, and may specifically include functional functions and parameters in the function set of the internal representation model of the deep learning compiler.

[0064] It should be noted that IR is an abstract level between high-level languages ​​and low-level machine codes. It provides an intermediate form independent of specific languages ​​and hardware, allowing compilers to better understand and process the structure and semantics of programs. Since IR is independent of specific languages ​​and hardware, it can be used for different programming languages ​​and platforms. Based on IR, compilers can convert source code from one language to another, or port code from one platform to another. By using IR, code porting and cross-compilation on different platforms can be achieved.

[0065] The input of the deep learning compiler is a quantized deep learning model in the format of one of the deep learning frameworks. The deep learning compiler can call the corresponding interface function according to the framework type of the deep learning model to convert the deep learning model into a unified computational graph. The output of the deep learning model is function and parameters. Among them, the function is the unified computational graph converted by the deep learning module, which can be saved in txt format; the parameter is the parameter data of the deep learning model, which can be saved in the format of a dictionary.

[0066] Exemplarily, the deep learning compiler can extract models from other frameworks (such as TensorFlow, PyTorch, ONNX). Next, the deep learning compiler translates the deep learning model into a high-level language model Relay, which is an intermediate expression file in an embodiment of the present invention. The model imported into the deep learning compiler is represented in Relay, which is a functional language and intermediate representation of the deep learning model. The deep learning model can be characterized by the functions and parameters in the function set of the deep learning compiler internal representation function.

[0067] Taking the TVM compiler as an example, conceptually, TVM can be divided into two layers: the Relay layer and the tir layer, which are run through the IRModule. In the embodiment of the present invention, the Relay layer is mainly used to convert the deep learning model into RelayIR. Specifically, the front-end component imports the deep learning model into the IRModule, which contains a set of functions that represent the model inside TVM. The compiler converts the IRModule into an IRModule that is functionally equivalent or approximately equivalent to it, that is, the intermediate expression file in the embodiment of the present invention.

[0068] For example, you can call the Python interface provided by TVM to convert the deep learning model into functions and parameters in TVM IR. Functions can be written into relay_function, and parameters can be written into relay_parameters. Relay_function corresponds to an end-to-end model, which can be simply understood as a computational graph that supports control flow, recursion, and complex data structures.

[0069] After the deep learning model to be compiled is converted into an intermediate expression file using a deep learning compiler, a target compiler matching the GPGPU is further obtained, which is used to translate the deep learning model into an equivalent target model in machine language format. Exemplarily, the target model can be an LLVM compiler. It can be understood that if GPGPU supports the framework of the model output by a compiler after compilation, that is, the model output by the compiler can be normally inferred and run on the GPGPU, the compiler can be considered to be a target compiler matching the GPGPU.

[0070] In an embodiment of the present invention, the target compiler is used to compile the intermediate expression file output by the deep learning compiler, and the operating environment corresponding to the deep learning model is constructed to obtain the compiled target model. It can be understood that the target model is a neural network model in a machine language format supported by GPGPU, and its function is equivalent to that of the deep learning model before compilation.

[0071] Taking the target compiler as LLVM as an example, the operating environment is also the runtime, which generally refers to the most basic software required for the code to run. Under normal circumstances, the runtime may also include a runtime library and a runtime system. Among them, the runtime library is a special computer program library used by the compiler to implement the built-in functions of the programming language to provide runtime (execution) support for the language program. This library generally contains basic input and output or memory management support. The runtime library is a library file required by the program at runtime, usually provided in the form of LIB or DLL. The runtime system refers to an environment in which the semi-compiled running code runs on the target machine. All languages ​​based on the runtime system are semi-compiled and semi-interpreted languages, such as C# and VisualBasic.NET under Java and .NET framework. The runtime system is an operating mode between compilation and interpretation. The compiler first compiles the source code into an intermediate code, and the runtime acts as an interpreter to interpret it during execution. In an embodiment of the present invention, the target compiler interprets the intermediate code of the deep learning model by constructing a runtime.

[0072] In an embodiment of the present invention, the functions and parameters contained in the intermediate expression file can be used to configure the runtime corresponding to the deep learning model, such as the library files and parameters required for the deep learning model to perform reasoning and operation on GPGPU.

[0073] Optionally, compiling the intermediate expression file using the target compiler and constructing a running environment corresponding to the deep learning model to obtain a compiled target model includes:

[0074] Step S11, configuring the graph structure, library and parameters of the deep learning model in the runtime construction script based on the intermediate expression file;

[0075] Step S12: execute the runtime build script to generate an operating environment for the deep learning model.

[0076] Exemplarily, the graph, library, and parameters of the deep learning model can be introduced into the runtime build script based on the intermediate expression file of the deep learning model. For example, the command line corresponding to the graph, library, and parameters of the deep learning model can be configured in the runtime build script according to the functions and parameters in the function set of the internal representation model of the deep learning compiler in the intermediate expression file. Then, by executing the runtime script, the runtime environment corresponding to the deep learning model can be generated, completing the conversion from the deep learning model to the target model.

[0077] After compiling the target model, the target image data to be processed can be imported into the target model for processing to obtain the target processing result, thereby realizing the compilation and operation of the deep learning model on the GPGPU platform.

[0078] The compilation method provided by the embodiment of the present invention can first call the deep learning compiler to convert the deep learning model into an intermediate expression file of the deep learning compiler, and then use the target compiler matching the GPGPU architecture to compile the intermediate expression file, and build a corresponding operating environment, so as to compile the deep learning model into a target model that can be run on the GPGPU, and use the target model to process the target image data on the GPGPU. The embodiment of the present invention does not need to consider the conversion problem of different model formats, nor does it need to configure the inference engine specifically for the GPGPU platform, which simplifies the model compilation process, reduces the compilation complexity, and improves the model compilation efficiency.

[0079] Optionally, the calling of the deep learning compiler in step 101 to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler includes:

[0080] Step S21, calling the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model;

[0081] Step S22: setting the input data type object and tensor shape of the deep learning model;

[0082] Step S23: configure functions and parameters corresponding to the deep learning model in the intermediate expression file of the deep learning compiler based on the large binary file, the data type, and the tensor shape.

[0083] When calling the deep learning compiler to convert the deep learning model, you can first download the deep learning model to be compiled and the corresponding binary large file (Blob) from the website or third-party platform (such as a server, a device used to build a deep learning model, etc.). It can be understood that the binary large file Blob is a container object containing binary data. The Blob object represents an immutable, raw data file-like object, and its data can be read in text or binary format.

[0084] After downloading the model and the large binary file, you can set the input data type (dtype) and tensor shape (shape) of the deep learning model, and then configure the corresponding functions and parameters of the deep learning model in the intermediate expression file of the deep learning compiler based on the large binary file, data type and tensor shape.

[0085] Specifically, you can use the corresponding interface in the deep learning compiler to download deep learning models and large binary files from websites or third-party platforms.

[0086] Then, call the interface or function in the deep learning model to set the input data type object and tensor shape of the deep learning model according to the data type that the deep learning model can process, the format requirements of the input data, and other information.

[0087] Finally, according to the large binary file of the deep learning model, the set input data type and tensor shape, the functions and parameters corresponding to the deep learning model are configured in the intermediate expression file.

[0088] Optionally, before importing the target image data to be processed into the target model for processing, the method further includes:

[0089] The image to be processed is preprocessed to convert the image to be processed into target image data; the target image data is text data in a target format.

[0090] In an embodiment of the present invention, before using the compiled target model for image processing, the image to be processed can be preprocessed to convert the image to be processed in a picture format into text data in a target format, that is, target image data, to facilitate parsing and reasoning by the target model.

[0091] For example, the image to be processed can be analyzed, and various image parameters can be filled into the image data function, where the image parameters can include: image storage path (imgpath), specified area (inshape), average value (mean), scaling value (scale), method (method), etc.

[0092] Optionally, the step of importing the target image data to be processed into the target model for processing to obtain a target processing result includes:

[0093] The target image data to be processed is imported into the target model, and the operating environment is called so that the target model can infer the target image data and output a target processing result.

[0094] In an embodiment of the present invention, after the target image data is imported into the target model, the operating environment (i.e., during the runtime) can be further called to run the target model to infer the target image data and output the target processing result. Exemplarily, this can be achieved by the following steps:

[0095] 1. Get input data, that is, target image data:

[0096] #get input

[0097] nd0 = nn_deploy.to_ndarray(data.astype('unint8'), tvm.ImgFormat.RGB,[inshape[2], inshape[3]])

[0098] 2. Set model input:

[0099] #set model input

[0100] module.set_input(nd0)

[0101] 3. Run the model:

[0102] #model run

[0103] model.run()

[0104] 4. Get the output data, that is, the target processing result:

[0105] #get output

[0106] nd_out = model.get_outputs()

[0107] In summary, the embodiment of the present invention provides a compilation method, which can first call the deep learning compiler to convert the deep learning model into an intermediate expression file of the deep learning compiler, and then use the target compiler matching the GPGPU architecture to compile the deep learning model, and use the intermediate expression file to build a corresponding operating environment, so as to compile the deep learning model into a target model that can run on the GPGPU, and use the target model to process the target image data on the GPGPU. The embodiment of the present invention does not need to consider the conversion problem of different model formats, nor does it need to configure the inference engine specifically for the GPGPU platform, which simplifies the model compilation process, reduces the compilation complexity, and improves the model compilation efficiency.

[0108] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0109] Device Embodiment

[0110] Reference Figure 2 , shows a structural block diagram of an embodiment of a compiling device of the present invention, the device is applied to a general-purpose graphics processor, and the device may specifically include:

[0111] The model conversion module 201 is used to call the deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler; the intermediate expression file includes functions and parameters in the function set of the internal representation model of the deep learning compiler;

[0112] A compiler acquisition module 202 is used to acquire a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format;

[0113] A model compilation module 203 is used to compile the intermediate expression file using the target compiler, and build an operating environment corresponding to the deep learning model to obtain a compiled target model;

[0114] The model inference module 204 is used to import the target image data to be processed into the target model for processing to obtain a target processing result.

[0115] Optionally, the model conversion module includes:

[0116] A download submodule, used to call the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model;

[0117] A setting submodule, used to set the input data type object and tensor shape of the deep learning model;

[0118] A configuration submodule is used to configure functions and parameters corresponding to the deep learning model in an intermediate expression file of the deep learning compiler based on the large binary file, the data type, and the tensor shape.

[0119] Optionally, the model compilation module includes:

[0120] A script configuration submodule, for configuring the graph structure, library and parameters of the deep learning model in a runtime build script based on the intermediate expression file;

[0121] The runtime generation submodule is used to execute the runtime construction script to generate the operating environment of the deep learning model.

[0122] Optionally, the device further comprises:

[0123] The preprocessing module is used to preprocess the image to be processed and convert the image to be processed into target image data; the target image data is text data in a target format.

[0124] Optionally, the model compilation module includes:

[0125] The compiling submodule is used to import the target image data to be processed into the target model and call the operating environment so that the target model can infer the target image data and output a target processing result.

[0126] In summary, the embodiment of the present invention provides a compilation device, which can first call the deep learning compiler to convert the deep learning model into an intermediate expression file of the deep learning compiler, and then use the target compiler matching the GPGPU architecture to compile the deep learning model, and use the intermediate expression file to build a corresponding operating environment, so as to compile the deep learning model into a target model that can run on the GPGPU, and use the target model to process the target image data on the GPGPU. The embodiment of the present invention does not need to consider the conversion problem of different model formats, nor does it need to configure the inference engine specifically for the GPGPU platform, which simplifies the model compilation process, reduces the compilation complexity, and improves the model compilation efficiency.

[0127] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0128] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0129] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0130] An embodiment of the present invention provides an electronic device for compiling, the electronic device is applied to a general-purpose graphics processor, the electronic device includes a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations:

[0131] Calling a deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler; the intermediate expression file includes functions and parameters in a function set of the internal representation model of the deep learning compiler;

[0132] Obtaining a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format;

[0133] Compile the intermediate expression file using the target compiler, and build an operating environment corresponding to the deep learning model to obtain a compiled target model;

[0134] The target image data to be processed is imported into the target model for processing to obtain a target processing result.

[0135] Optionally, calling a deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler includes:

[0136] Calling the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model;

[0137] Set the input data type object and tensor shape of the deep learning model;

[0138] Based on the large binary file, the data type, and the tensor shape, functions and parameters corresponding to the deep learning model are configured in the intermediate expression file of the deep learning compiler.

[0139] Optionally, compiling the intermediate expression file using the target compiler and constructing a running environment corresponding to the deep learning model to obtain a compiled target model includes:

[0140] Configuring the graph structure, libraries and parameters of the deep learning model in a runtime build script based on the intermediate expression file;

[0141] Execute the runtime build script to generate an operating environment for the deep learning model.

[0142] Optionally, the electronic device is further configured to execute, by one or more processors, the one or more programs including instructions for performing the following operations:

[0143] The image to be processed is preprocessed to convert the image to be processed into target image data; the target image data is text data in a target format.

[0144] Optionally, the step of importing the target image data to be processed into the target model for processing to obtain a target processing result includes:

[0145] The target image data to be processed is imported into the target model, and the operating environment is called so that the target model can infer the target image data and output a target processing result.

[0146] Figure 3 1 is a block diagram of an electronic device 600 for compiling according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0147] Reference Figure 3 , the electronic device 600 may include one or more of the following components: a processing component 602 , a memory 604 , a power component 606 , a multimedia component 608 , an audio component 610 , an input / output (I / O) interface 612 , a sensor component 614 , and a communication component 616 .

[0148] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.

[0149] The memory 604 is configured to store various types of data to support operations on the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0150] The power supply component 606 provides power to the various components of the electronic device 600. The power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 600.

[0151] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0152] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), and when the electronic device 600 is in an operation mode, such as a call mode, a recording mode, and a voice information processing mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 604 or sent via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.

[0153] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0154] The sensor assembly 614 includes one or more sensors for providing various aspects of status assessment for the electronic device 600. For example, the sensor assembly 614 can detect the open / closed state of the electronic device 600, the relative positioning of the components, such as the display and keypad of the device 600, and the sensor assembly 614 can also detect the position change of the electronic device 600 or a component of the electronic device 600, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and the temperature change of the electronic device 600. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0155] The communication component 616 is configured to facilitate wired or wireless communication between the electronic device 600 and other devices. The electronic device 600 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency information processing (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0156] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0157] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the instructions can be executed by a processor 620 of an electronic device 600 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0158] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device (server or terminal), enables the processor to perform Figure 1 The compilation method shown.

[0159] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0160] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, devices or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0161] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems) and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0162] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0164] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0165] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0166] The above is a detailed introduction to a compilation method, device and electronic device provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A compiling method, characterized in that: Applied to a general-purpose graphics processor, the method comprises: Calling a deep learning compiler to extract a deep learning model from other frameworks, and calling a corresponding interface function according to the framework type of the deep learning model to convert the deep learning model into an intermediate expression file; the intermediate expression file includes functions and parameters in a function set of the internal representation model of the deep learning compiler; Obtaining a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format; Compiling the intermediate expression file using the target compiler and building an operating environment corresponding to the deep learning model to obtain a compiled target model; the operating environment includes library files and parameters required for the deep learning model to perform reasoning and operation on the general-purpose graphics processor; The target image data to be processed is imported into the target model for processing to obtain a target processing result.

2. The method according to claim 1, characterized in that The calling of the deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler includes: Calling the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model; Set the input data type object and tensor shape of the deep learning model; Based on the large binary file, the data type, and the tensor shape, functions and parameters corresponding to the deep learning model are configured in the intermediate expression file of the deep learning compiler.

3. The method according to claim 1, characterized in that The using the target compiler to compile the intermediate expression file and construct an operating environment corresponding to the deep learning model to obtain a compiled target model includes: Configuring the graph structure, libraries and parameters of the deep learning model in a runtime build script based on the intermediate expression file; Execute the runtime build script to generate an operating environment for the deep learning model.

4. The method according to claim 1, characterized in that: Before importing the target image data to be processed into the target model for processing, the method further includes: The image to be processed is preprocessed to convert the image to be processed into target image data; the target image data is text data in a target format.

5. The method according to claim 1, characterized in that The step of importing the target image data to be processed into the target model for processing to obtain a target processing result includes: The target image data to be processed is imported into the target model, and the operating environment is called so that the target model can infer the target image data and output a target processing result.

6. A compiling device, characterized in that: The device is applied to a general-purpose graphics processor, and comprises: A model conversion module, used to call a deep learning compiler to convert the deep learning model to be compiled into an intermediate expression file of the deep learning compiler; the intermediate expression file includes functions and parameters in a function set of the internal representation model of the deep learning compiler; A compiler acquisition module, used to acquire a target compiler that matches the architecture of the general-purpose graphics processor; the target compiler is used to translate the deep learning model into an equivalent target model in machine language format; A model compilation module, used to compile the intermediate expression file using the target compiler, and build an operating environment corresponding to the deep learning model to obtain a compiled target model; the operating environment includes library files and parameters required for the deep learning model to perform reasoning and operation on the general-purpose graphics processor; A model reasoning module is used to import the target image data to be processed into the target model for processing to obtain a target processing result; Wherein, the model conversion module is specifically used for: Call the deep learning compiler to extract the deep learning model from other frameworks, and call the corresponding interface function according to the framework type of the deep learning model to convert the deep learning model into an intermediate expression file.

7. The device according to claim 6, characterized in that The model conversion module comprises: A download submodule, used to call the deep learning compiler to download the deep learning model to be compiled and the large binary file corresponding to the deep learning model; A setting submodule, used to set the input data type object and tensor shape of the deep learning model; A configuration submodule is used to configure functions and parameters corresponding to the deep learning model in an intermediate expression file of the deep learning compiler based on the large binary file, the data type, and the tensor shape.

8. The device according to claim 6, characterized in that The model compilation module includes: A script configuration submodule, for configuring the graph structure, library and parameters of the deep learning model in a runtime build script based on the intermediate expression file; The runtime generation submodule is used to execute the runtime construction script to generate the operating environment of the deep learning model.

9. An electronic device, characterized in that: The electronic device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to execute the compiling method according to any one of claims 1 to 5 by one or more processors.

10. A readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the processor is enabled to execute the compilation method as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Hardware adaptation device and method based on deep learning

    CN114186678A