Neural Network Compilation Method, System, Computer Device, and Storage Medium

By combining the compilation process with the verification process, generating verification models and providing software comparison results and hardware comparison files, the problem of long development cycle of CNN hardware accelerator is solved, and a more efficient development and testing process is achieved.

CN114399019BActive Publication Date: 2025-06-10SHANGHAI XINZHENG ELECTRONICS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111647960.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-10
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

In the prior art, the development cycle of CNN hardware accelerator is long, mainly due to the separation of the test verification process and the compilation process, making it difficult to locate the error position. Iterative error correction requires redeployment of hardware, and the verification process cycle is long.

Method used

A neural network compilation method is proposed to generate a verification model by analyzing standard model data and quantizing data, and compiling and generating hardware configuration files in combination with hardware parameters. This method combines the compilation process with the verification process, and provides software comparison results and hardware comparison files through the generated verification model, positioning and solving pre-compilation and deployment errors.

Benefits of technology

By combining the compilation process with the verification process, the test and development cycle in the product development stage is shortened, development efficiency is improved, and the number of hardware iterative deployments is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399019B_ABST
    Figure CN114399019B_ABST
Patent Text Reader

Abstract

The present application discloses a neural network compilation method, system, computer device, and storage medium. The method includes inputting the standardized model data and quantization data of the network to be compiled, as well as the hardware parameters of the hardware platform to be deployed; first performing parsing processing; automatically generating a verification model according to the parsing result; converting the parsed data into network model configuration data for compilation; and combining with the hardware parameters to obtain a calculation loop chunking scheme; determining the current compilation mode, if it is the test mode, then running the generated verification model to obtain a software comparison result and generating a hardware comparable file; finally, compiling and generating a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop chunking scheme. The neural network compilation method combines the compilation process with the verification process, so that in the product R & D stage, through the generated verification model, a software comparison result and a hardware comparable file are provided, accelerating the test R & D speed and shortening the development cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of neural networks, and particularly to a neural network compilation method, system, computer device, and storage medium. Background Art

[0002] A field programmable gate array (FPGA) is a semi-custom circuit with strong reconfigurability and low implementation latency, and is used as a hardware platform for accelerating the inference calculation of convolutional neural networks (CNNs). However, with the expansion of the application fields of CNNs, more and more new networks with more complex hierarchical structures have been proposed. Some deep CNN models have dozens to hundreds of layers, and there are significant differences in size and configuration between layers. These changing trends in the CNN model structure increase the complexity of hardware design, making it more difficult to design a general CNN hardware accelerator to effectively map different CNN algorithms. As the scale and complexity of CNN models continue to increase, the method of customized design becomes less and less practical, and the method of automated compilation is essential.

[0003] The existing compilation method for deploying CNN models on the FPGA platform is based on the CNN model structure. The calculations of different types of layers in the network model are mapped to the computing units of the hardware accelerator to be deployed, and a hardware configuration file that can control the behavior of the accelerator is compiled and generated. After the compilation is completed, the generated hardware configuration file is deployed offline on the hardware accelerator, and the implementation result of the CNN model on the hardware can be obtained.

[0004] The above compilation method controls the hardware behavior through the hardware configuration file. By flexibly changing different configuration information in the hardware configuration file, different mapping implementations of the CNN model on the hardware accelerator can be obtained. To verify the accuracy of the configuration information in the compiled hardware configuration file, the existing method is to verify through the hardware inference result of the CNN hardware accelerator. Specifically, it is usually necessary to manually build a Python model or other software models of the network to be tested, generate software inference results and then compare them with the hardware inference results obtained by deployment. However, the above verification method is separated from the compilation process, and it is difficult to locate the specific error position in the specific implementation process. In the subsequent iterative error correction process, the hardware needs to be redeployed, and the verification process has a long cycle, which further leads to a long development cycle of the CNN hardware accelerator. Summary of the Invention

[0005] To solve the problem that in the existing compilation method, the test verification process is separated from the compilation process, the relevance is not strong, it is difficult to locate the error position, and the hardware needs to be redeployed during the iterative error correction process, and the verification process has a long cycle, which further leads to a long development cycle of the CNN hardware accelerator, this application provides a neural network compilation method, system, computer device and storage medium in the following aspects.

[0006] In the first aspect of this application, a neural network compilation method is provided, including:

[0007] Input the standard model data and quantization data of the neural network to be compiled, as well as the hardware parameters of the hardware platform to be deployed;

[0008] Parse the standard model data and quantization data and perform software inference operations to obtain intermediate model structure information and intermediate parameter data; among them, the intermediate model structure information includes model structure description information and model structure parsing data;

[0009] Automatically generate a corresponding verification model according to the model structure description information;

[0010] Preprocess the intermediate model structure information and intermediate parameter data to obtain network model configuration data; among them, the network model configuration data includes a compilable network hierarchy structure information, the weight parameters of the neural network to be compiled, the input test data of the neural network to be compiled, and the software inference result;

[0011] According to the network model configuration data, combined with the hardware parameters, obtain a calculation loop cutting scheme for the neural network to be compiled;

[0012] Judge whether the current execution mode is the test mode or the user mode;

[0013] If it is the test mode, then perform the following operations:

[0014] Run the verification model, obtain the verification model calculation result, compare it with the software inference result, and obtain the software comparison result;

[0015] Generate a hardware comparable file according to the verification model calculation result;

[0016] Compile and generate a hardware configuration file according to the network model configuration data, hardware parameters and calculation loop cutting scheme;

[0017] If it is the user mode, then perform the following operations:

[0018] Compile and generate a hardware configuration file according to the network model configuration data, hardware parameters and calculation loop cutting scheme.

[0019] Optionally, automatically generating a corresponding verification model according to the model structure description information includes:

[0020] Input the model structure description information, and create a calculation code document and a call code document for the verification model;

[0021] Traverse all layers in the model structure description information, and layer by layer write the corresponding calculation operation codes into the calculation code document according to the layer type; among them, the calculation operations in the corresponding calculation operation codes are implemented by running the calculation functions of the corresponding operations in the preset verification model function library;

[0022] Traverse all layers in the model structure description information, and write the processing operation codes of the calculation results of the corresponding layers into the call code document according to the layer type in the model structure description information;

[0023] Write the corresponding codes for generating software comparison result files and hardware comparable files into the call code document of the verification model;

[0024] Obtain the verification model.

[0025] Optionally, run the verification model, obtain the verification model calculation results, compare them with the software inference results, and obtain the software comparison results, including:

[0026] Run the call code document and load the software inference results;

[0027] Input the input test data and hardware parameters of the neural network to be compiled, run the calculation code of the target layer in the verification model for calculation, obtain the calculation results of the target layer in the verification model, and compare the data of the corresponding layer in the software inference results according to the calculation results of the target layer in the verification model to obtain the software comparison results of the target layer; where the target layer is any layer in the verification model;

[0028] Traverse all layers in the verification model to obtain the verification model calculation results and the software comparison results.

[0029] The second aspect of this application provides a neural network compilation system for implementing the steps of a neural network compilation method provided in the first aspect of this application; the neural network system includes:

[0030] A model parsing module, which is used to parse the standard model data and quantization data and perform software inference operations to obtain intermediate model structure information and intermediate parameter data; among them, the intermediate model structure information includes model structure description information and model structure parsing data;

[0031] The model conversion module is used to perform the following operations: preprocess the intermediate model structure information and intermediate parameter data to obtain network model configuration data; wherein, the network model configuration data includes compilable network hierarchy information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results; according to the network model configuration data, combined with the hardware parameters of the hardware platform to be deployed, obtain a calculation loop slicing scheme for the neural network to be compiled.

[0032] The judgment module is used to judge whether the current execution mode of the compilation system is the user mode or the test mode.

[0033] The model compilation module is used to compile and generate a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop slicing scheme.

[0034] The model verification module is used to perform the following operations: automatically generate a corresponding verification model according to the model structure description information; run the verification model to obtain the calculation result of the verification model, compare it with the software inference result to obtain a software comparison result; generate a hardware comparable file according to the calculation result of the verification model.

[0035] Optionally, the model verification module is further used to perform the following operations:

[0036] Input the model structure description information, and create a calculation code document and a call code document for the verification model.

[0037] Traverse all layers in the model structure description information, and layer by layer write the corresponding calculation operation code into the calculation code document according to the layer type; wherein the calculation operation in the corresponding calculation operation code is implemented by running the calculation function of the corresponding operation in the preset verification model function library.

[0038] Traverse all layers in the model structure description information, and write the processing operation code of the calculation result of the corresponding layer into the call code document according to the layer type in the model structure description information.

[0039] Write the corresponding code for generating the software comparison result file and the hardware comparable file into the call code document of the verification model.

[0040] Obtain the verification model.

[0041] Optionally, the model verification module is further used to perform the following operations:

[0042] Run the call code document and load the software inference result.

[0043] Input the input test data and hardware parameters of the neural network to be compiled, run the calculation code of the target layer in the verification model for calculation, obtain the calculation result of the target layer in the verification model, and compare the data of the corresponding layer in the software inference result according to the calculation result of the target layer in the verification model to obtain the software comparison result of the target layer; where the target layer is any layer in the verification model.

[0044] Traverse all layers in the verification model to obtain the verification model calculation result and the software comparison result.

[0045] The third aspect of the present application discloses a computer device, including:

[0046] A memory for storing a computer program;

[0047] A processor for implementing the steps of the neural network compilation method disclosed in the first aspect of the present application when executing the computer program.

[0048] The fourth aspect of the present application discloses a computer-readable storage medium, and the storage medium stores a computer program, and when the computer program is processed and executed, it implements the steps of the neural network compilation method disclosed in the first aspect of the present application.

[0049] The present application discloses a neural network compilation method, system, computer device and storage medium through the above aspects. The method includes inputting the standardized model data and quantization data of the network to be compiled, as well as the hardware parameters of the hardware platform to be deployed; parsing the standardized model data and quantization data to obtain intermediate model structure information including model structure description information and intermediate parameter data; automatically generating a verification model according to the model structure description information; preprocessing the intermediate model structure information and intermediate parameter data to obtain network model configuration data that can be used for hardware compilation; combining the hardware parameters to obtain a calculation loop slicing scheme for the neural network to be compiled; judging the current compilation mode, if it is the test mode, then run the generated verification model to obtain the verification model calculation result, compare it with the software inference result in the network model configuration data to obtain the software comparison result, and generate a hardware comparable file; finally, compile and generate a hardware configuration file according to the network model configuration data, hardware parameters and the calculation loop slicing scheme; if it is the user mode, directly compile and generate a hardware configuration file according to the network model configuration data, hardware parameters and the calculation loop slicing scheme. The neural network compilation method combines the compilation process with the verification process, provides software comparison results and hardware comparable files through the generated verification model, speeds up the test and research speed in the product research and development stage, and shortens the development cycle.

[0050] Furthermore, the verification model in the neural network compilation method disclosed in this embodiment can simulate the hardware inference behavior, so that error problems can be solved at the software level before compilation and deployment. At the same time, the iterative error correction and redeployment at the hardware level are transformed into verification error correction at the software level, greatly shortening the test development cycle of the product. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic diagram of the working process of a neural network compilation method disclosed in an embodiment of the present application;

[0052] Figure 2 is a schematic diagram of the working process of generating a verification model in a neural network compilation method disclosed in an embodiment of the present application;

[0053] Figure 3 is a schematic diagram of the working process of running a verification model in a neural network compilation method disclosed in an embodiment of the present application;

[0054] Figure 4 is a schematic diagram of the structure of a neural network compilation system disclosed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] In order to solve the problems that the test verification process and the compilation process in the existing compilation method are separated, the relevance is not strong, it is difficult to locate the error position, and the hardware needs to be redeployed during the iterative error correction process, and the verification process has a long cycle, which further leads to a long development cycle of the CNN hardware accelerator, the present application provides a neural network compilation method, system, computer device and storage medium through the following embodiments.

[0056] See Figure 1 , a neural network compilation method disclosed in an embodiment of the present application. The method realizes an automatic compilation process according to the standard model data and quantization data of the neural network to be compiled and the hardware parameters of the hardware platform to be deployed, generates a configuration file required by the hardware to be deployed, and combines the corresponding verification model to generate a software comparison result and a hardware comparable file, shortening the development cycle and accelerating product iteration. As Figure 1 shown, the neural network compilation method includes the following steps:

[0057] Step 10, input the standard model data and quantization data of the neural network to be compiled, and the hardware parameters of the hardware platform to be deployed.

[0058] In this embodiment, the standard model data uses an ONNX (Open Neural Network Exchange) standardized model file. ONNX is an open format for representing deep neural network models and is used to store trained models. ONNX defines a set of standard formats independent of the environment and platform, providing a basis for the interoperability of deep learning models, enabling deep learning models to be used interchangeably in different frameworks and environments. The ONNX file stores not only the weights of the neural network model but also the model's structural information, the input and output of each layer in the network, and some other auxiliary information. Quantization refers to converting the floating-point algorithm of the neural network to be compiled into a fixed-point algorithm. In this embodiment, the quantization data includes quantization processing information for different operations in different layers. The hardware parameters of the hardware platform to be deployed include, but are not limited to, data bit width, computing parallelism, and on-chip storage resources.

[0059] The neural network to be compiled described in this application includes, but is not limited to, CNN, DNN (Deep Neural Networks), RNN (Recurrent Neural Network), LSTM (Long short-term memory), SNN (Spiking Neuron Networks), and Transformer models.

[0060] Step 20: Parse the standard model data and quantization data and perform software inference operations to obtain intermediate model structure information and intermediate parameter data.

[0061] In this embodiment, an analysis tool is used to first analyze the ONNX model data of the neural network to be compiled and the corresponding quantization data, and then further software inference operations are performed based on the analyzed data. The analysis process refers to extracting the model's structural information and network parameters by parsing the deep learning model description file, and performing operations such as connecting Tensor transformation, operator merging, and operator splitting on the model computation graph according to the hardware's constraint conditions. Combining with the inference operation of the neural network, the specific software results of each layer can be obtained. Through the above operations, intermediate model structure information and intermediate parameter data can be obtained. Among them, the hardware's constraint conditions include computing resource constraints and storage resource constraints.

[0062] In practical applications, the intermediate model structure information includes model structure description information (prototxt document) and model structure parsing data; the intermediate parameter data includes parameter data related to network inference such as weight quantization bias, network input data provided by a random generation program or a document loading program, and software layer-by-layer inference result data. Among them, the model structure description information is a special format document with the file suffix ".prototxt", which is called a prototxt document in this application. The prototxt document is used to store information describing the neural structure of a neural network and can be viewed using the caffe (Convolutional Architecture for Fast Feature Embedding, a convolutional neural network framework) tool.

[0063] In some examples, an ONNX parsing tool written in the python language is used to parse the ONNX model data and the corresponding quantization data. Other languages such as caffe can also be used to write the above ONNX parsing tool.

[0064] Step 30: Automatically generate a corresponding verification model according to the model structure description information.

[0065] In this embodiment, a golden verification model is used as the verification model in the neural network compilation method. The generation result of the golden verification model is stored in the form of matrix elements and can be directly used as the software comparison result to be compared and verified element by element with the software inference result. At the same time, since the golden verification model is a verification design based on the hardware architecture, the generated calculation result can also be flexibly converted into a hardware comparable format. If there is a problem when comparing with the software inference result, the error location can be quickly located and the problem can be solved by quickly iterating and verifying the update. This verification model can solve the error problem at the software level one step before compilation and deployment, and convert the cumbersome deployment error correction operation at the hardware level into a fast verification and error correction operation at the software level. In this embodiment, matlab is used to implement the generation of the golden verification model to make the verification model better fit the calculation process of the hardware platform. In other embodiments, other programming languages such as C language can also be used to implement the verification model.

[0066] Further, referring to Figure 2 , step 30, automatically generating a corresponding verification model according to the model structure description information includes:

[0067] Step 31: Input the model structure description information and create a calculation code document and a call code document for the verification model.

[0068] In this embodiment, the prototxt document obtained in input step 20 is used as the basic document for generating the golden verification model, and a verification model calculation code document corresponding to the neural network to be compiled and a corresponding call code document are newly created.

[0069] Step 32: Traverse all layers in the model structure description information, and layer by layer write the corresponding calculation operation code into the calculation code document according to the layer type; wherein the calculation operation in the corresponding calculation operation code is implemented by running the calculation function of the corresponding operation in the preset verification model function library.

[0070] Step 33: Traverse all layers in the model structure description information, and write the processing operation code of the calculation result of the corresponding layer into the call code document according to the type of the layer in the model structure description information.

[0071] Step 34: Write the corresponding code for generating the software comparison result file and the hardware comparable file into the call code document of the verification model.

[0072] Step 35: Obtain the verification model corresponding to the neural network to be compiled.

[0073] In this embodiment, according to the network structure information in the prototxt document, a golden verification model corresponding to the network to be compiled is automatically generated. The golden verification model includes a verification model calculation code document and a corresponding call code document.

[0074] In practical applications, first load the prototxt document of the neural network to be compiled for reading operation, and newly create a calculation code document of the corresponding golden verification model for writing operation. Traverse all layers in the prototxt document, judge the type of the current layer, and automatically write the corresponding calculation operation code into the corresponding position in the calculation code document according to the layer type to obtain the calculation code of the corresponding layer in the verification model. Among them, the calculation operation in the corresponding calculation operation code is implemented by running the calculation function of the corresponding operation in the preset verification model function library. The types of layers include but are not limited to batch normalization, ReLU, pooling, element addition calculation, deconvolution. The preset verification model function library includes pre-written calculation functions related to the verification model calculation operation. According to the types corresponding to different layers of the model, the calculation functions include but are not limited to convolution calculation functions, pooling layer calculation functions, ReLU operation calculation functions, etc. After writing the corresponding verification model calculation code layer by layer in the calculation code document, the calculation code document of the verification model is obtained.

[0075] In practical applications, after completing the writing of the calculation code document for the verification model, open the call code document (golden_top document) of the verification model for writing operations, and automatically write the code for the preprocessing part of the golden_top document. Traverse all layers in the prototxt document, perform cumulative counting operations on different types of layers, and determine whether there is a certain type of layer based on whether the cumulative result is greater than 0. Automatically write the code for processing the calculation results of the above-mentioned type of layer into the golden_top document; among them, the types of layers include, but are not limited to, batch normalization, ReLU, pooling, element-wise addition calculation, and deconvolution; the code for processing the calculation results of the above-mentioned type of layer is used to perform operations including saving data results and printing data results. And automatically write the corresponding code for generating software comparison result files and hardware comparable files into the call code document. After completion, obtain the call code document golden_top document of the verification model. Subsequently, automatically load the input test data of the neural network to be compiled and run the verification model by calling the golden_top document, and output the corresponding software comparison result files and hardware comparable files.

[0076] It should be noted that in practical applications, as long as step 30 and the corresponding steps 31-36 are executed after step 20 and before step 70, the same technical effects of this embodiment can be achieved.

[0077] Step 40: Preprocess the intermediate model structure information and intermediate parameter data to obtain network model configuration data; among them, the network model configuration data includes compilable network hierarchy information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results.

[0078] In practical applications, since hardware processing is a regular and neat operation, it is necessary to preprocess the model structure and corresponding parameter data obtained by the parsing tool to meet the requirements of subsequent calculations. In step 40 of this embodiment, the above-mentioned intermediate model structure information and intermediate parameter data are converted into network model configuration data that can be recognized by the hardware compiler and adapted to subsequent processing. Specifically, the network model configuration data includes compilable network hierarchy information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results. The above-mentioned network model configuration data can be directly used for hardware compilation processing, and some of the parameters can also be used as the input of the verification model to generate the test comparison results of the golden verification model.

[0079] Step 50: According to the network model configuration data and combined with the hardware parameters of the hardware platform to be deployed, obtain a calculation loop slicing scheme for the neural network to be compiled.

[0080] To better deploy the neural network to be compiled onto the corresponding hardware platform and make full use of on-chip resources, it is necessary to perform slicing processing on the dimensions of the network model. The network model after slicing processing can better adapt to the storage and computing resources of the hardware accelerator to be deployed, so as to maximize the on-chip resource utilization rate and data reuse, and minimize data communication. Among them, the computational loop tiling scheme is a commonly used hardware optimization and acceleration scheme, which cuts the network structure of the model along certain dimensions of the computational loop to relieve the pressure of data transmission inside and outside the chip. At the same time, when the on-chip resources are limited, the on-chip resources can be utilized more effectively.

[0081] Step 60, determine whether the current execution mode is the test mode or the user mode.

[0082] The neural network compilation method disclosed in this embodiment provides two working modes, the test mode and the user mode. The test mode is used during the development and testing of the corresponding compilation system to ensure the correctness of the model parsing and compilation process of the network, so as to accelerate the development cycle. The user mode is used by the user during the use of the corresponding compilation system, and only needs to parse and compile the neural network model to be compiled to generate a hardware configuration file. During actual use, when running the corresponding compilation system, the compilation requirements need to be input. In some implementation manners, the compilation requirements include whether the current mode of running the compilation system is the user mode or the test mode, whether it is necessary to compare with the golden verification model calculation result, whether it is necessary to parse the batch normalization layer, whether it is necessary to print the on-board test result, whether the model needs to be automatically sliced, whether it is necessary to load the input test data document of the neural network to be compiled or randomly generate the input data of the neural network to be compiled, whether it is per-channel quantization, etc.

[0083] If it is the test mode, then execute steps 70 to 90.

[0084] Step 70, run the verification model, obtain the verification model calculation result, compare it with the software inference result, and obtain the software comparison result.

[0085] In this embodiment, according to the current compilation mode, if it is the test mode, the automatically generated golden verification model is run, and after a series of calculations, the verification model calculation result is obtained. The obtained verification model calculation result is compared with the software inference result obtained through preprocessing in step 40 to obtain the software comparison result.

[0086] Further, refer to Figure 3 , step 70, run the verification model, obtain the verification model calculation result, compare it with the software inference result, and obtain the software comparison result, including:

[0087] Step 71, run the calling code document of the verification model and load the software reasoning result.

[0088] In this embodiment, the code golden_top document corresponding to the main calling function of the verification model is called to run the golden verification model, and at the same time, step 40 is loaded to obtain the software reasoning result as a basis for subsequent software comparison.

[0089] Step 72, input the input test data of the neural network to be compiled and the hardware parameters of the hardware platform to be deployed, run the calculation code of the target layer in the verification model to perform calculations, obtain the calculation results of the target layer in the verification model, and compare the data of the corresponding layer in the software reasoning result according to the calculation results of the target layer in the verification model to obtain the software comparison results of the target layer, where the target layer is any layer in the verification model.

[0090] Input the input test data of the neural network to be compiled in the network model configuration data obtained in step 40, and input the hardware parameters of the hardware platform to be deployed. Configure the relevant parameters in the golden verification model according to the hardware parameters. According to the network structure of the neural network to be compiled, call the calculation code layer by layer to generate the verification model calculation results.

[0091] Because the calculation results and software reasoning results of the verification model are stored in the form of matrix arrays and can be directly compared, the calculation results of the target layer in the verification model and the software reasoning results of the corresponding layer can be compared layer by layer and element by element to obtain the software comparison results of the target layer.

[0092] Step 73, traverse all layers in the verification model to obtain verification model calculation results and software comparison results.

[0093] In the actual application process, the obtained software comparison results of each layer are written into the software comparison result document, and the above software comparison results are used as a reference for the rapid verification and error correction operation of the corresponding software layer.

[0094] Step 80, generating a hardware comparable file based on the calculation results of the verification model.

[0095] In this embodiment, the calculation code of the corresponding function of the hardware comparable file is run using the call code document golden_top document of the verification model, and the hardware comparable data for each layer is generated according to the calculation results of the verification model for each layer. The hardware comparable data for each layer is written into the hardware comparable file to obtain a hardware comparable file that can be compared layer by layer with the hardware implementation results. The operations for generating the hardware comparable file mainly include changes in the dimensions of the data matrix, conversion of the data matrix in the calculation results of the verification model into the form stored in the hardware, data base conversion, and conversion from decimal to binary and hexadecimal data. Among them, the optionally generated hardware comparable file includes a simulation comparison data file and a board comparison data file.

[0096] Step 90, compile and generate a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop chunking scheme.

[0097] In this embodiment, the network model configuration data includes compilable network hierarchy information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results. The hardware parameters include, but are not limited to, data bit width, calculation parallelism, and on-chip storage resources. Based on the network model configuration data, hardware parameters, and calculation loop chunking scheme, data transmission configuration (including DMA scheduling instructions, etc.), overall scheduling configuration, internal configuration of module units (including register parameter assignment, etc.), and storage configuration of the hardware platform are performed to generate a hardware configuration file. The hardware platform configures the relevant registers according to the generated hardware configuration file as required by the compilation, and stores the calculation data, parameters, and scheduling instructions in the specified arrangement order.

[0098] If it is the user mode, directly execute step 90.

[0099] Step 90, compile and generate a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop chunking scheme.

[0100] In this embodiment, in the user mode, without model verification, directly execute step 90 to obtain a hardware configuration file. In practical applications, whether it is the test mode or the user mode, ultimately a hardware configuration file needs to be generated to complete the compilation work of the network model for deploying the network model to the corresponding hardware platform.

[0101] The compiled and generated hardware configuration file can control the behavior of the hardware accelerator by changing the register configuration or hardware parameters; at the same time, the verification model in the test mode can locate the problem location through the generated software comparison results, and solve problems such as operator forced adaptation and parsing errors that can only be discovered by the existing compiler through on-board operation.

[0102] This embodiment discloses a neural network compilation method, including inputting the standardized model data and quantization data of the network to be compiled, as well as the hardware parameters of the hardware platform to be deployed; parsing the standardized model data and quantization data to obtain intermediate model structure information including model structure description information and intermediate parameter data; automatically generating a verification model according to the model structure description information; preprocessing the intermediate model structure information and intermediate parameter data to obtain network model configuration data that can be used for hardware compilation; combining the hardware parameters to obtain a calculation loop slicing scheme for the neural network to be compiled; judging the current compilation mode, if it is the test mode, running the generated verification model to obtain the calculation results of the verification model, comparing them with the software inference results in the network model configuration data to obtain a software comparison result, and generating a hardware comparable file; finally, compiling and generating a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop slicing scheme; if it is the user mode, directly compiling and generating a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop slicing scheme. The neural network compilation method combines the compilation process with the verification process, provides a software comparison result and a hardware comparable file through the generated verification model, ensures the correctness of the network model parsing and compilation process, speeds up the test and research speed in the product R & D stage, and shortens the development cycle.

[0103] Furthermore, the verification model in the neural network compilation method disclosed in this embodiment can simulate the hardware inference behavior, so that error problems can be solved at the software level one step before compilation and deployment. At the same time, converting the iterative error correction and redeployment at the hardware level into verification and error correction at the software level greatly shortens the test and development cycle of the product.

[0104] The second embodiment of this application provides a neural network compilation system for implementing the steps of a neural network compilation method disclosed in the first embodiment. Refer to Figure 2 , the neural network compilation system provided in this embodiment includes a model parsing module, a model conversion module, a judgment module, a model compilation module, and a verification module.

[0105] The model parsing module is used to parse the standard model data and quantization data and perform software inference operations to obtain intermediate model structure information and intermediate parameter data; among them, the intermediate model structure information includes model structure description information and model structure parsing data.

[0106] The model conversion module is used to perform the following operations: preprocess the intermediate model structure information and intermediate parameter data to obtain network model configuration data; wherein, the network model configuration data includes compilable network hierarchy information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results; according to the network model configuration data and in combination with the hardware parameters of the hardware platform to be deployed, obtain a calculation loop slicing scheme for the neural network to be compiled.

[0107] The judgment module is used to judge whether the current execution mode of the compilation system is the user mode or the test mode.

[0108] The model compilation module is used to compile and generate a hardware configuration file according to the network model configuration data, hardware parameters, and calculation loop slicing scheme.

[0109] The model verification module is used to perform the following operations: automatically generate a corresponding verification model according to the model structure description information; run the verification model to obtain the calculation result of the verification model, compare it with the software inference result to obtain a software comparison result; generate a hardware comparable file according to the calculation result of the verification model.

[0110] Furthermore, when the model verification module is used to automatically generate a corresponding verification model, it is used to perform the following operations:

[0111] Step 31, input the model structure description information, and create a calculation code document and a call code document for the verification model.

[0112] Step 32, traverse all layers in the model structure description information, and layer by layer write the corresponding calculation operation code into the calculation code document according to the layer type; wherein the calculation operation in the corresponding calculation operation code is implemented by running the calculation function of the corresponding operation in the preset verification model function library.

[0113] Step 33, traverse all layers in the model structure description information, and according to the layer type in the model structure description information, write the processing operation code of the calculation result of the corresponding layer into the call code document.

[0114] Step 34, write the corresponding code for generating the software comparison result file and the hardware comparable file into the call code document of the verification model.

[0115] Step 35, obtain the verification model.

[0116] Furthermore, the model verification module is also used to perform the following operations:

[0117] Step 71, run the call code document of the verification model and load the software inference result.

[0118] Step 72: Input the input test data of the neural network to be compiled and the hardware parameters of the hardware platform to be deployed, run the calculation code of the target layer in the verification model for calculation, obtain the calculation result of the target layer in the verification model, and compare the data of the corresponding layer in the software inference result according to the calculation result of the target layer in the verification model to obtain the software comparison result of the target layer; where the target layer is any layer in the verification model.

[0119] Step 73: Traverse all layers in the verification model to obtain the calculation result of the verification model and the software comparison result.

[0120] This embodiment gives two examples for the generation process and verification process of the model verification module. The neural network to be compiled described in this embodiment includes, but is not limited to, CNN, DNN, RNN, LSTM, SNN, Transformer models. The following examples take two CNN models as the neural network models to be compiled to illustrate the execution process of this embodiment.

[0121] Example 1: Take the resnet18 network model with a quantization data bit width of 16bit as the model to be compiled to illustrate the generation process of the model verification module and its executed functions. Load the parsed prototxt document of the resnet18 network model for reading, and create a calculation code document and a call code document for the golden verification model. First, open the calculation code document for writing. Determine the type of the current layer. If it is a convolution operation, write the code for the convolution calculation operation in the verification model calculation code document, where the convolution calculation function in the preset verification model function library is called in the calculation code. In the resnet18 model, the calculation types of the network model include convolution, max pooling, average pooling, element addition, fully connected layer, etc. Traverse all layers of the resnet18 network to complete the writing of the calculation code document for the golden verification model.

[0122] Then open the call code document of the golden verification model (golden_top document), perform write operations according to the prototxt document of the network model, and automatically write the code in the preprocessing part of the golden_top document. Traverse all layers in the prototxt document, perform cumulative counting operations on different types of layers, and judge whether there is a certain type of layer according to whether the cumulative result is greater than 0. Automatically write the processing operation code corresponding to the calculation result of the above type of layer into the golden_top document; among them, the types of layers include but are not limited to batch normalization, ReLU, pooling, element addition calculation, and deconvolution; and automatically write the corresponding code for generating software comparison result files and hardware comparable files (including simulation comparison data files and on-board comparison data files). After completion of writing, obtain the code golden_top document corresponding to the main call function of the verification model. Subsequently, by calling the golden_top document, automatically load the input test data of the neural network to be compiled and run the verification model, and output the corresponding software comparison result files and hardware comparable files.

[0123] When the test mode is enabled, call the generated golden_top document, load the input test data of the neural network to be compiled and the corresponding comparison data obtained by software inference in the network model configuration data obtained in the data preprocessing part of the model conversion module, as well as the hardware parameters of the platform to be deployed. Run the calculation code in the calculation code document of the golden verification model of the network to be verified, and perform operations layer by layer according to the resnet18 network model structure. Among them, the input test data of the verification model should be consistent with the input test data in the software inference process in the model parsing module. During the calculation process, when specific operations such as convolution, pooling, and batch normalization are involved, call the corresponding calculation functions in the preset golden verification model function library. For each layer calculated to obtain the calculation result of that layer in the verification model, compare it with the data result of that layer in the loaded software inference data until all layers in all networks are calculated, and print the software comparison result files and hardware comparison files of each layer.

[0124] Example 2: Take the VGG16 network model with a quantization data bit width of 16bit as the model to be compiled to illustrate the generation process of the model verification module and the functions it performs. First, load the parsed prototxt document of the VGG16 network for read operations, open the corresponding calculation code document of the golden verification model for write operations, traverse all layers and write the corresponding code to complete the writing of the calculation code document of the golden verification model; then open the corresponding call code document of the golden verification model for write operations, and write the corresponding processing operation code according to the layer types existing in the VGG16 network to complete the call code document of the golden verification model; obtain the verification model corresponding to the VGG16 network.

[0125] When the test mode is enabled, the calling code document corresponding to the VGG16 network verification model loads the network model configuration data and hardware parameters and runs the verification model, calls different computing functions to calculate layer by layer and compare the calculation results layer by layer. After all layers are calculated, the software comparison result file and the hardware comparison file of each layer are printed.

[0126] In practical applications, the neural network compilation system can also be referred to as a neural network compilation tool chain to vividly represent a series of processing processes implemented by the neural network compilation system.

[0127] Through the above content, the second embodiment of the present application discloses a neural network compilation system for implementing the steps of a neural network compilation method disclosed in the first embodiment. The system includes two working modes: user mode and test mode. The user mode is used when the user turns it on, and the test mode is used during the development and testing of the compilation system to verify the correctness of the model parsing and compilation process. The neural network compilation system provided in this embodiment combines the automated verification process with the network compilation process, provides software comparison results and hardware comparable files through the verification model generated by the model verification module, speeds up the test and research and development speed in the product research and development stage, and shortens the development cycle.

[0128] It should be noted that before running the compilation system, in addition to inputting the standard model data and quantization data of the neural network to be compiled and the hardware parameters of the hardware platform to be deployed, it is also necessary to configure the current compilation requirements according to the actual situation, such as whether the current mode of running the compilation system is user mode or test mode, whether it is necessary to compare the verification results, whether it is necessary to parse the batch normalization layer, whether it is necessary to print the on-board test results, whether the model needs to be automatically sliced, whether it is necessary to load test data or randomly generate data, and whether it is channel-by-channel quantization, etc.

[0129] The third embodiment of the present application discloses a computer device, including a memory and a processor; wherein the memory is used to store a computer program; the processor is used to implement the steps of the neural network compilation method as described in the first embodiment of the present application when executing the computer program.

[0130] The fourth embodiment of the present application discloses a computer-readable storage medium. The storage medium stores a computer program, and the computer program realizes the steps of the neural network compilation method as described in the first embodiment of the present application when being processed and executed.

[0131] The present application has been described in detail above in conjunction with specific embodiments and exemplary examples, but these descriptions should not be construed as limiting the present application. Those skilled in the art understand that, without departing from the spirit and scope of the present application, various equivalent substitutions, modifications, or improvements can be made to the technical solutions and their implementation manners of the present application, and all of these fall within the scope of the present application. The protection scope of the present application shall be subject to the appended claims.

[0132] For the similarities and identical parts among the above various embodiments, reference can be made to each other.

Claims

1. A neural network compilation method, characterized in that, it includes: Input the standard model data and quantization data of the neural network to be compiled, as well as the hardware parameters of the hardware platform to be deployed; Parse the standard model data and quantization data and perform software inference operations to obtain intermediate model structure information and intermediate parameter data; among them, the intermediate model structure information includes model structure description information and model structure parsing data; Automatically generate a corresponding verification model according to the model structure description information, including: Input the model structure description information, and create a calculation code document and a call code document for the verification model; Traverse all layers in the model structure description information, and layer by layer write the corresponding calculation operation code into the calculation code document according to the layer type; among them, the calculation operation in the corresponding calculation operation code is implemented by running the calculation function of the corresponding operation in the preset verification model function library; Traverse all layers in the model structure description information, and write the processing operation code of the calculation result of the corresponding layer into the call code document according to the layer type in the model structure description information; Write the corresponding code for generating a software comparison result file and a hardware comparable file in the call code document; Obtain the verification model; Preprocess the intermediate model structure information and the intermediate parameter data to obtain network model configuration data; among them, the network model configuration data includes compilable network hierarchy information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results; According to the network model configuration data, combined with the hardware parameters, obtain a calculation loop slicing scheme for the neural network to be compiled; Judge whether the current execution mode is a test mode or a user mode; If it is the test mode, then perform the following operations: Run the verification model to obtain the verification model calculation result, compare it with the software inference result to obtain a software comparison result; Generate a hardware comparable file according to the verification model calculation result; Compile and generate a hardware configuration file according to the network model configuration data, the hardware parameters, and the calculation loop slicing scheme; If it is the user mode, then perform the following operations: Compile and generate a hardware configuration file according to the network model configuration data, the hardware parameters, and the calculation loop slicing scheme.

2. The neural network compilation method according to claim 1, characterized in that, The running of the verification model to obtain the verification model calculation result, comparing it with the software inference result to obtain a software comparison result includes: Run the call code document and load the software inference result; Input the input test data of the neural network to be compiled and the hardware parameters, run the calculation code of the target layer in the verification model for calculation, obtain the calculation result of the target layer in the verification model, and compare the data of the corresponding layer in the software inference result according to the calculation result of the target layer in the verification model to obtain the software comparison result of the target layer; where the target layer is any layer in the verification model; Traverse all layers in the verification model to obtain the calculation result of the verification model and the software comparison result.

3. A neural network compilation system Characterized in that It is used to implement the steps of a neural network compilation method described in claim 1 or 2; the neural network compilation system includes: A model parsing module, which is used to parse standard model data and quantization data and perform software inference operations to obtain intermediate model structure information and intermediate parameter data; among them, the intermediate model structure information includes model structure description information and model structure parsing data; A model conversion module, which is used to perform the following operations: preprocess the intermediate model structure information and the intermediate parameter data to obtain network model configuration data; among them, the network model configuration data includes a compilable network hierarchy structure information, weight parameters of the neural network to be compiled, input test data of the neural network to be compiled, and software inference results; according to the network model configuration data, combined with the hardware parameters of the hardware platform to be deployed, obtain a calculation loop slicing scheme for the neural network to be compiled; A judgment module, which is used to judge whether the current execution mode of the compilation system is the user mode or the test mode; A model compilation module, which is used to compile and generate a hardware configuration file according to the network model configuration data, the hardware parameters, and the calculation loop slicing scheme; A model verification module, which is used to perform the following operations: automatically generate a corresponding verification model according to the model structure description information; The model verification module is further used to perform the following operations: Input the model structure description information, and create a calculation code document and a call code document for the verification model; Traverse all layers in the model structure description information, and layer by layer write the corresponding calculation operation code into the calculation code document according to the layer type; among them, the calculation operation in the corresponding calculation operation code is implemented by running the calculation function of the corresponding operation in a preset verification model function library; Traverse all layers in the model structure description information, and write the processing operation code of the calculation result of the corresponding layer into the call code document according to the layer type in the model structure description information; Write the corresponding code for generating a software comparison result file and a hardware comparable file into the call code document; Obtain the verification model; Run the verification model to obtain the verification model calculation result, compare it with the software inference result to obtain the software comparison result; generate a hardware comparable file according to the verification model calculation result.

4. A neural network compilation system according to claim 3 Characterized in that The model verification module is further used to perform the following operations: Run the call code document and load the software inference result; Input the input test data of the neural network to be compiled and the hardware parameters, run the calculation code of the target layer in the verification model for calculation, obtain the calculation result of the target layer in the verification model, and compare the data of the corresponding layer in the software inference result according to the calculation result of the target layer in the verification model to obtain the software comparison result of the target layer; wherein the target layer is any layer in the verification model. Traverse all layers in the verification model to obtain the calculation result of the verification model and the software comparison result.

5. A computer device Characterized in that It includes: A memory for storing computer programs; A processor for implementing the steps of the neural network compilation method as described in claim 1 or 2 when executing the computer program.

6. A computer-readable storage medium Characterized in that The storage medium stores a computer program, and when the computer program is processed and executed, it implements the steps of the neural network compilation method as described in claim 1 or 2.

Citation Information

Patent Citations

  • Automated design method, device and optimization method applied for neural network processor

    CN107016175A

  • Method, system and device for compiling AI chip and medium

    CN112232497A