A Method for Automatically Compiling and Running C / C++ Code for Huawei Ascend Accelerator Cards
Through the unified description model and memory management model of custom operator functions, combined with automated configuration programs, the custom C/C++ code development process on Astend Accelerator Card is simplified, the complicated steps of custom code for Astend Accelerator Card is solved, and efficient custom operator development and call are realized.
Patent Information
- Application Number
- CN202111533075.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Ascend Accelerator Card has complicated steps when executing custom C/C++ code, and it is difficult for the existing technology to develop specific high-performance operators efficiently and conveniently.
The unified description model of custom operator functions, the data model of the host and Ascend accelerated card processor memory management, the automatic configuration program of the Ascend platform custom operator, and the call execution system of the custom operator are formed to form a software framework to simplify the development and call process of the custom operator.
It realizes efficient execution of custom function code on Ascend Acceleration Card, simplifies development steps, improves development efficiency and convenience of code operation.
Smart Images

Figure CN114461186B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method for automatically compiling and running C / C++ code for Huawei Ascend acceleration cards. Background Art
[0002] Ascend acceleration cards are high-performance and low-power AI acceleration modules developed by Huawei. They provide extremely powerful computing capabilities and offer multi-level programming interfaces based on CANN (heterogeneous computing architecture for AI scenarios) for building AI applications. Currently, Ascend acceleration cards are mainly used in the construction, training, and inference of AI models that require large computing power. Recently, the demands for parallel computing applications and high-performance computing applications have been increasing day by day. Both of these types of applications also require significant computing power support. As a representative of domestic high-computing-power machines, Ascend acceleration cards are also an option for running these two types of applications in the future.
[0003] Ascend acceleration cards provide the unified programming language AscendCL to the upper layer. The operators in the C language API library provided by AscendCL are all existing general-purpose operators, which are available for building model networks but not suitable for parallel high-performance applications. Custom operators can be implemented by registering one's own functions with AscendCL, thereby using Ascend to build more specific high-performance applications.
[0004] AscendCL supports custom AI CPU operators in C / C++ language. Logic execution code and configuration files need to be written according to specific patterns, including operator model definition, operator model format, operator implementation, etc., and are compiled by a specific compiler and finally copied to a specific directory of an external operator library. The programs that can run on Ascend acceleration cards highly depend on the existing operator libraries and operation models. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, when using Ascend acceleration cards, the steps for executing custom code are numerous and cumbersome. The present invention aims to provide a method for automatically compiling and running C / C++ code for Huawei Ascend acceleration cards. Specifically, based on the custom AI CPU operators that can be defined by the Ascend unified programming language AscendCL, a more general C / C++ operator development framework is proposed, with the goal of enabling more efficient and convenient development of specific high-performance operators and building AI applications and high-performance applications with special needs on the Ascend platform.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for automatically compiling and running C / C++ code for Huawei Ascend acceleration cards. By leveraging the Ascend platform provided by the Ascend acceleration card processor to the upper layer and the characteristics of its custom operators, combined with C / C++ language compilation, the overall process of operator function development and invocation is integrated with the data management and operation scheduling capabilities of the host and the Ascend acceleration card processor to achieve the execution of custom function code on the Ascend acceleration card. This method is ultimately implemented in the form of a software framework. It includes a unified description model for custom operator functions, a data model for memory management of the host and the Ascend acceleration card processor, an automated configuration program for custom operators on the Ascend platform, and a call execution system for custom operators.
[0008] It should be noted that the unified description model of the custom operator function is to adapt to custom functions with various types of parameters on the premise of meeting the constraints of custom operators specified by the Ascend platform. Through template programming and the method of redundant parameters, the unified model of the operator function can cover various types of data and standardized input and output parameters.
[0009] It should be noted that the data model for memory management of the host and the Ascend acceleration card processor integrates the memory allocation methods of the main memory and the Ascend acceleration card processor, and the data copy between the two. Through the data description structure of the data model, a json description file is generated during operator invocation, and then the operator is invoked.
[0010] It should be noted that the automated configuration program for custom operators on the Ascend platform is an automated module that generates a configuration file based on the custom operators defined by the user, the invoked operators, and the data model instances, and simultaneously processes the logic code of the Ascend platform specifications that is irrelevant to the operators.
[0011] It should be noted that the call execution system for custom operators is a module that automatically processes the call logic of custom functions. Based on the data model and the automated configuration program, the call execution system also strengthens the management of the executed functions and devices. It uses a hash table to map the functions to be executed by the user and the Ascend acceleration card device on which they are executed. It only provides a run function for the user to execute each custom function, simplifying the call process of custom operators in the Ascend platform. Brief Description of the Drawings
[0012] Figure 1 It is a schematic diagram of the development and call process of general custom AI CPU operators;
[0013] Figure 2 It is a schematic diagram of developing and calling custom operators using the framework of the present invention;
[0014] Figure 3This is a schematic diagram of the overall process of the framework of the present invention for processing custom operators;
[0015] Figure 4 This is a schematic diagram of the development and compilation process of the operator processed by the framework of the present invention;
[0016] Figure 5 This is a schematic diagram of the operator call process processed by the framework of the present invention. Detailed implementation manners
[0017] The present invention will be further described below in conjunction with the accompanying drawings. It should be noted that the following embodiments are based on the present technical solution and give detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.
[0018] The present invention is a method for automatically compiling and running C / C++ code for Huawei Ascend acceleration cards. By using the Ascend platform provided by the Ascend acceleration card processor to the upper layer and the characteristics of its custom operators, combined with C / C++ language compilation, the overall process of operator function development and call is integrated, and the data management and operation scheduling capabilities of the host and the Ascend acceleration card processor are integrated to realize the execution of custom function code on the Ascend acceleration card. This method is finally implemented in the form of a software framework; among them, it includes a unified description model of custom operator functions, a data model for memory management of the host and the Ascend acceleration card processor, an automated configuration program for custom operators on the Ascend platform, and a call execution system for custom operators.
[0019] It should be noted that the unified description model of the custom operator function is to enable the framework of the present invention to adapt to custom functions with various types of parameters on the premise of meeting the constraints of the custom operators specified by the Ascend platform; through template programming and the method of redundant parameters, the unified model of the operator function can cover various types of data and standardized input and return values.
[0020] It should be noted that the data model for memory management of the host and the Ascend acceleration card processor integrates the memory allocation methods of the main memory and the memory allocation of the Ascend acceleration card processor, and through the data copy between the two, a json description file is generated by the data description structure of the data model during operator call, and then the operator is called.
[0021] It should be noted that the automated configuration program for custom operators on the Ascend platform is an automated module that generates a configuration file according to the custom operators defined by the user, the called operators, and the data model instances, and at the same time processes the logic code of the Ascend platform specifications that has nothing to do with the operators.
[0022] It should be noted that the call execution system of the custom operator is a module that automatically processes the custom function call logic. Based on the data model and the automated configuration program, the call execution system also strengthens the management of execution functions and devices, uses a hash table to map the functions to be executed by the user, and on which Ascend acceleration card device to execute; only provides a run function for the user, and each custom function can be executed, simplifying the custom operator call process for the Ascend platform.
[0023] Embodiment
[0024] First, a brief description of the abbreviations and key terms used in this embodiment is given below:
[0025] CANN: Unified heterogeneous computing framework, by providing multi-level programming interfaces, supporting users to quickly build AI applications and services based on the Ascend platform.
[0026] AscendCL (Ascend Computing Language): A set of C language API libraries for developing applications on the Ascend platform, mainly used to manage operations, call existing AI models and operator operations, so as to realize the computing power of the Ascend acceleration card.
[0027] Custom operator: In the present invention, it refers to the operator operations not provided by AscendCL or the operators already existing in the operator library that one wants to re-implement oneself. Generally speaking, it is to run the C / C++ functions written by oneself on the Ascend acceleration card.
[0028] To make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the present invention will be described in more detail below in conjunction with the accompanying drawings in the present invention.
[0029] Figure 1 Describes the general development and call process of custom AI CPU operators, and this process strictly conforms to the steps and content required by the AscendCL specification. It is mainly divided into two parts: operator development and operator call, and is associated by the operator name (name) through the operator call function provided by AscendCL.
[0030] The operator development in the example involves operator prototype definition, operator parameter data format definition and operator implementation, and an operator adaptation plugin based on a third-party framework (such as pytorch, tensorflow) can also be written.
[0031] The code for operator development needs to be compiled using a specific code library and compiler aarch-gnu to generate the dynamic link library file of the operator. Finally, the header file, configuration file and dynamic link library file of the operator need to be copied to a specific directory of the AscendCL external operator library.
[0032] The operator calls in the example are divided into three steps:
[0033] 1) Resource preparation, including device resources and operator resources. The application for device resources is stipulated by the usage mode of AscendCL, and the operator resources are the operator model and operator data. Both are necessary parts for running the operator.
[0034] 2) Run the operator. Before running, the data in the main memory needs to be transferred to the memory of the acceleration card device, and then the operator is run using the operator name string. After the operation is completed, the result will be output to the specified memory area.
[0035] 3) Release resources. After the entire model operation is completed, the occupied resources need to be released.
[0036] The entire custom operator has many steps from development to invocation, and there are also several configuration files, including the operator definition configuration file, model file, data description file, etc. There are also multiple management operations stipulated by AscendCL. These do not involve the operator logic itself. Automated management is a more appropriate way, allowing developers to focus on the development of the operator's logic code rather than platform adaptation.
[0037] Figure 2 It is a schematic diagram for developing and invoking a custom operator function using the framework of the present invention. Only the operator code needs to be developed, and the operator can be directly invoked after compilation.
[0038] When developing a custom operator using the framework, only the logic code of the operator itself needs to be concerned about. On the code framework defined by the framework, fill in the processing logic of the operator and execute the compilation instruction, then the custom operator function can be invoked.
[0039] The framework automates the non-code logic configuration files that were originally required during operator development, shortening the processing flow. Developers only need to develop the most core C / C++ logic code of the operator.
[0040] After the development of the operator's logic code is completed, only a simple compilation command needs to be executed for compilation, and then the operator function can be used.
[0041] The framework automates most of the resource preparation work during the invocation period, allowing developers to invoke the operator function without awareness. Only by using the run function provided by the framework can the corresponding operator be directly invoked.
[0042] Comparison Figure 1The framework simplifies the entire custom operator usage process, automates most of the configurations, manages device resources without the need for developers to pay extra attention, and proposes a data model for efficient management of memory resources. Using the framework significantly reduces the steps of using custom operators and improves development efficiency.
[0043] Figure 3 , Figure 4 and Figure 5 This is a schematic diagram of the processing flow of the core framework of the present invention, which shows the overall processing content of the framework and the processing content performed during the development and call of the processing operator.
[0044] Figure 3 This is a schematic diagram of the overall process of processing custom operators in the framework of the present invention. The main steps are as follows:
[0045] 1) First, the framework needs to integrate the pure C / C++ code written by the developer, without the user having to rewrite other configuration files.
[0046] 2) The framework itself is configured with the corresponding compilation dependency libraries and specific compilers, and compilation can be completed through simple cmake and make commands.
[0047] 3) Automatic registration of custom operators during compilation.
[0048] 4) During the operator calling phase, the framework provides an initialization function to automatically apply for device resources.
[0049] 5) The framework provides a data model to manage data, memory distribution and data copying.
[0050] 6) The framework provides a run function to execute operators, which can automatically execute the corresponding custom operators.
[0051] 7) After the calculation is completed, the framework provides a finalize function to automatically recycle device resources.
[0052] Figure 4 Schematic diagram of the compilation process of the framework processing operator development of the present invention.
[0053] In order to better handle the pure C / C++ code customized by developers, the framework has designed a unified operator model. The unified operator model supports operator functions with inconsistent numbers of parameters by providing multiple function interfaces with different numbers of parameters, and implements dynamic parameter type determination at compile time through template metaprogramming. The combination of the above two solves the problem of customized operators with different numbers of parameters and different types of parameters, and supports the development of multiple different customized operator functions at the same time.
[0054] The operator unified model processes the operator prototype description file and the operator configuration file through an automated configuration program, integrates multiple operator codes into the same configuration file, and differentiates different operators only according to the parameter type and quantity. This avoids generating too many configuration files and no longer requires users to pay attention to configuration files unrelated to the code.
[0055] The framework supports compiling Ascend acceleration card executable files by configuring the necessary dynamic link libraries and specific compilers. Just execute the cmake and make compilation instructions, which greatly reduces the user's usage cost.
[0056] To meet the requirements of AscendCL, after compilation, the framework's automated configuration program automatically copies the compiled dynamic link library files and operator configuration files to the external operator library directory to achieve the registration of custom operators on AscendCL.
[0057] Figure 5 This is a schematic diagram of the process of the framework of the present invention for processing operator calls.
[0058] During the use of the operator, four stages of processes are required:
[0059] 1) Create a running environment. Before using the Ascend acceleration card for computing, it is necessary to first apply for available device resources, and the framework provides an init function to initialize the resources. The operator model file (.om) and the operator call data structure description file (.json) are generated through the data description structure of the data model to initialize the device and the computing context.
[0060] 2) Data management and memory space management. This stage is the stage of preparing and initializing operator data. The data model provided by the framework supports memory applications in different spaces, quickly applies for memory spaces on the host and the device, and supports freely deep-copying data at different memory locations.
[0061] 3) Execute operator computing. After the data preparation is completed, the run function of the framework is responsible for executing the operator call and performing the operation on the acceleration card device. The data model is passed into the run function, and the run function passes the data memory address and data description structure of the data model to the operator. When executing, the corresponding dynamic link library (.so) is called according to the operator name of the operator unified model, and the corresponding specific custom operator function is executed according to the parameters. After the operator is executed, the calculation result will be saved in the device memory and can be used for further calculation, or the result can be copied back to the main memory through the data model.
[0062] 4) Resource release. The framework provides a finalize function to automatically release the occupied resources, including Device, Context, etc.
[0063] The following is the development of a 3D convolution operator using the framework of the present invention.
[0064] 1) Analyze according to the input and output parameters of the 3D convolution operator. This operator has four input parameters, namely three-dimensional data and its dimensional information, three-dimensional kernel data and its dimensional information, and the output is three-dimensional data and its dimensional information. The data used in the 3D convolution calculation process is of float type. Therefore, select the operator function template with 4 input and 4 output parameters in the unified operator model to fill the logic code.
[0065] 2) Write the logic code of the 3D convolution operator function. The code written using the operator function template and the code written independently with the same function are basically the same. Therefore, it is relatively easy to change the original function to the function code of the unified operator model.
[0066] 3) In the preparation for calling, construct a data description structure according to the data of the 3D convolution and initialize the device.
[0067] 4) Prepare the data, instantiate the data model, specify the allocated memory space, and perform necessary data transfers, mainly to transfer to the accelerator card device.
[0068] 5) Call the run function to execute the 3D convolution operator calculation.
[0069] 6) After the calculation is completed, copy the calculation result back to the main memory and compare it with the calculation result of the general 3D convolution. The results are consistent.
[0070] 7) After all the calculations are completed, release the occupied resources.
[0071] For those skilled in the art, various corresponding changes and deformations can be given according to the above technical solutions and concepts, and all these changes and deformations should be included within the protection scope of the claims of the present invention.
Claims
1. A method for automatically compiling and running C / C++ code for Huawei Ascend acceleration cards. By leveraging the characteristics of the Ascend platform and its custom operators provided by the Ascend acceleration card processor to the upper layer, combined with C / C++ language compilation, through the overall process of operator function development and invocation, and integrating the data management and operation scheduling capabilities of the host and the Ascend acceleration card processor, it realizes the execution of custom function code on the Ascend acceleration card. This method is finally implemented in the form of a software framework, and is characterized in that The method includes a unified description model for custom operator functions, a data model for memory management of the host and Ascend accelerator card processors, an automated configuration program for custom operators on the Ascend platform, and a call execution system for custom operators; the unified description model for custom operator functions is to adapt to custom functions with various types of parameters while meeting the constraints of custom operators specified by the Ascend platform; through template programming and the method of redundant parameters, the unified model of operator functions can cover various types of data and standardized input and return values; the data model for memory management of the host and Ascend accelerator card processors integrates the memory allocation methods of the main memory and the memory allocation of the Ascend accelerator card processors, and the data copy between the two. A json description file is generated through the data description structure of the data model during operator call, and then the operator is called; the automated configuration program for custom operators on the Ascend platform is an automated module that generates a configuration file according to the custom operators defined by the user, the called operators, and the data model instances, and at the same time processes the logic code that complies with the Ascend platform specifications and has nothing to do with the operators; the call execution system for custom operators is a module that automatically processes the call logic of custom functions. Based on the data model and the automated configuration program, the call execution system also strengthens the management of the execution functions and devices, uses a hash table to map the functions to be executed by the user, and on which Ascend accelerator card device to execute; only provides a run function for the user, and various custom functions can be executed, simplifying the call process of custom operators called on the Ascend platform.
Citation Information
Patent Citations
Reactor neutron transport calculation method based on domestic accelerator card
CN110704106A
Mercuric chloride processor management and scheduling method based on an SLURM job scheduling system
CN112882828A