Method and system for automatically generating reasoning codes in model deployment process
By automatically generating the model inference code on the HC3080 chip and quickly obtaining API parameters using operator information and index tables, the time-consuming generation of model inference code on the HC3080 chip is solved, efficient deployment and error self-testing are achieved, and model deployment efficiency of domestic chips is improved.
Patent Information
- Application Number
- CN202510430401.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing technology cannot efficiently simplify the generation of model inference codes on AI chips such as HC3080, especially complex models such as YOLOV7, which are time-consuming and lack methods to effectively improve the efficiency of inference code implementation.
By extracting model information from the intermediate file of the model to be deployed, defining operator variables and generating API calls, using operator names and index tables to quickly obtain API key parameters, and automatically generate inference codes based on the characteristics of the HC3080 chip, and including error self-testing function.
It significantly reduces the model deployment time, shortens from several days to within a few minutes, improves the model inference efficiency on the HC3080 chip, and ensures code accuracy through error self-test, enriching the ecosystem of domestic chips.
Smart Images

Figure CN120297418A_ABST
Abstract
Description
Technical Field
[0001] This solution belongs to the technical field of model deployment, and particularly relates to a method and system for automatically generating inference code during the model deployment process. Background Art
[0002] The deployment of AI models includes model quantization and model inference. Currently, the mainstream AI model quantization deployment tools in the market include TensorRT of NVIDIA, RKNN-Toolkit of Ruixin Microelectronics, and openppl of SenseTime, etc. Their common feature is that they will transform the quantized model into a form customized by their toolchains, thereby simplifying the model deployment process through this method.
[0003] However, for AI chips that do not receive the support of these tools, it is impossible to use these existing quantization deployment tools to simplify the development process and reduce the deployment cost. For model quantization, the applicant has proposed an automatic quantization deployment method and system for AI chips that support operator-level operations [Publication No.: CN118605890B], providing an automatic quantization and deployment method for AI chips such as HC3080 that are not supported by the above tools, which can improve the model deployment efficiency of such chips. For model inference, for AI chips such as HC3080 that are not supported by the aforementioned tools, to successfully infer the model on the AI chip, the traditional approach is for developers to implement it layer by layer using quantization information and the operator APIs provided by the AI chip, which is extremely time-consuming. Especially for complex models like YOLOV7 with more than two hundred operators, the deployment process takes about two or three working days. Currently, there is a lack of effective means to significantly improve the implementation efficiency of inference code. Summary of the Invention
[0004] The purpose of this solution is to propose a method and system for automatically generating inference code during the model deployment process in view of the problems existing in the prior art.
[0005] To achieve the above objective, this solution adopts the following technical solutions: A method for automatically generating inference code during the model deployment process, the method includes: Extracting model information from the intermediate file of the model to be deployed; the model information includes operator names, operator connection relationships, operator pad values, and the original sizes of each variable of the operator; Defining variables for each operator, including defining variable names according to operator names and defining memory allocation for variables according to the original sizes of each variable, so that the variables of the operator can be quickly located according to the operator name; the principle of memory allocation is to allocate memory for the corresponding variables according to the original size multiplied by a coefficient greater than or equal to 1, such as a coefficient of 1.2; The variables of an operator include output variables and weight variables; therefore, it includes defining the names of output variables and weight variables. Since the input variable is the output variable of the previous operator, the input variable is not redefined. Completing the output variable means completing the input variable of the next operator; Load the model data of the model to be deployed and generate operator weight loading code for each operator; Generate an output variable name index table and a weight variable name index table according to the order of operators in the intermediate file and the definition of variables for each operator; Determine the input variable name of each operator according to the connection relationship of operators, and generate an input variable name index table in the order of operators; Generate a pad value index table, an input size index table, and an output size index table during actual inference according to the association relationship of operators in the intermediate file and the model information; Two directly or indirectly connected operators are considered to have an association relationship; For each operator, extract the corresponding API from the API library, and fill the parameters of the API according to the output variable name index table, weight variable name index table, input size index table, output size index table, input variable name index table, and pad value index table to generate the API call of the corresponding operator; Combine the API calls of each operator to obtain the required inference code.
[0006] In the method for automatically generating inference code in the above model deployment process, the operator weight loading code for each operator is generated in the following manner: Define the weight data name of the corresponding operator according to the operator name, and the weight variable to be assigned can be quickly determined according to the operator name; Generate a first bin function for the operator to read the size of the weight data; Generate a second bin function for the operator to read the weight data and assign it to the weight variable, that is, store the read weight data into the corresponding memory allocated to the weight variable.
[0007] In the method for automatically generating inference code in the above model deployment process, defining variable names and weight data names according to operator names specifically includes: For the output variable, add a suffix indicating "output" at the end of the corresponding operator name as the output variable name; For the weight variable, add a suffix indicating "weight" at the end of the corresponding operator name as the weight variable name; For the weight data, add a suffix indicating "weight data" at the end of the corresponding operator name as the weight data name.
[0008] In the above method for automatically generating inference code in the model deployment process, variable names of each operator are defined in the form of pointers to define the memory allocation of variables while defining the variable names. The generated API calls also include quantization information, fixed parameters, and operator parameters obtained from intermediate files. The quantization information is obtained during the quantization process and will not be elaborated here. The operator parameters are determined according to specific operators and APIs.
[0009] In the above method for automatically generating inference code in the model deployment process, the generation of inference code also includes header file inclusion and general variable definition. After generating the inference code, this method further includes closing the file handle used to write the inference code into the C file.
[0010] In the above method for automatically generating inference code in the model deployment process, generating the pad value index table, input size index table, and output size index table during actual inference according to the association relationship of operators and model information in the intermediate file specifically includes: determining the pad value of the current operator based on the model information, and generating the pad value index table according to the order of operators in the intermediate file. The pad value is used to describe whether subsequent level operators associated with the current operator need padding and the size of the padding number. According to the pad values of each operator in the pad value index table and the output information of each operator in the intermediate file, that is, the original output size, the output size during inference of each operator is obtained to get the output size index table. According to the connection relationship of each operator in the intermediate file, the input size during inference of each operator is obtained to get the input size index table.
[0011] In the above method for automatically generating inference code in the model deployment process, the pad value of the current operator is determined in the following way: Extract the operator feature list of the intermediate file. The operator feature list is used to describe all operator features of the model to be deployed currently, where each element is an operator list of the corresponding operator. The operator list is used to describe the current operator, including the operator name, the pad value P_INDEX of the operator in the intermediate file, and the index of subsequent operators associated with the current operator. Traverse the operators after the current operator according to the operator feature list and the operator list, and take the maximum P_INDEX among them as the pad value of the current operator until the convolution operator, pooling operator, or the last operator of the model is traversed.
[0012] In the above method for automatically generating inference code in the model deployment process, this method further includes: Add inference error prediction code containing code macro switches to each of the aforementioned API calls, which is used to calculate the absolute value of the difference between the inference value and the theoretical value, and close the corresponding API call when the absolute value of the difference is greater than or equal to the preset difference value.
[0013] In the method for automatically generating inference code in the above model deployment process, the intermediate file is an ONNX structure model; This method is used to automatically generate inference code for the model to be deployed in the HC3080 chip.
[0014] An inference code automatic generation system for the model deployment process is used to automatically generate inference code for the model to be deployed by the above method.
[0015] The advantages of this solution are as follows: 1. This solution provides a method for automatically generating inference code for AI chips such as HC3080 that are not supported by existing quantization deployment tools, simplifying the model deployment process. For example, the deployment time of a model with more than two hundred operators like YOLOV7 can be reduced from two or three working days to a few minutes or even within one minute. 2. This solution proposes to automatically generate AI model inference code by obtaining the key parameters of the operator API, and the key parameters required for API calls can be quickly obtained through the operator name or index. 3. This solution proposes a method for recursively searching for the maximum pad value required by subsequent operators in intermediate files such as the onnx structure. Through this pad value and the characteristics of convolving first and then padding on specific AI chips such as HC3080, the same effect as padding first and then convolving in onnx can be achieved, greatly improving the model inference efficiency on such AI chips as HC3080. 4. This solution proposes an inference error self - checking function while realizing the automatic generation of inference code, which can calculate the inference error of each operator and give a prompt, ensuring the accuracy of the automatically generated inference code. 5. This solution is applicable to chips such as HC3080, enriching the ecosystem of such domestic chips. Description of the Drawings
[0016] Figure 1 It is the method flow chart of the method for automatically generating inference code in the model deployment process proposed by this solution; Figure 2 It is a schematic diagram of the definition method of weight variable names in the embodiment of this solution; Figure 3 It is a schematic diagram of the definition method of input variable names in the embodiment of this solution; Figure 4 It is the processing sequence diagram of the padding operation of the convolution operator during ONNX model inference; Figure 5 Schematic diagram of the Pad implementation of the convolution operator on the HC3080 chip in the embodiment of this solution; Figure 6 Code example diagram of the conv operator call and corresponding inference error calculation generated in the embodiment of this solution. Specific implementation manner
[0017] As Figure 1 shown, this solution provides a method for automatically generating inference code in the model deployment process and a system that can be used to execute this method. It can realize one-key generation of the inference code of the AI model on the AI chip, such as C code. In the model deployment process, it is usually necessary to convert the model to be deployed into an intermediate format. Currently, the widely used intermediate representation format is mainly the ONNX (Open Neural Network Exchange) format, which allows sharing of models between different deep learning frameworks. In this embodiment, the intermediate format is ONNX as an example. When put into use, it can also be replaced with other feasible intermediate formats, which are not specifically limited here.
[0018] This embodiment takes the typical HC3080 chip as an example to introduce the process of automatically generating inference code in detail. The principle of code automatic generation is very simple, which is to create a file handle and write the code string to be generated into a file with the suffix ".c". Therefore, how to obtain the code string is the most critical issue. This solution mainly divides the C code of model inference into five parts, including header file introduction and general variable definition, operator input / output and weight variable definition, loading input files and weight files, inference code generation, and file handle release.
[0019] Header file introduction and general variable definition The purpose of this embodiment is to generate the C code of the AI model that can be inferred on the HC3080. The beginning of the C language must include relevant header files and define general variables. Therefore, this embodiment first completes the automatic generation of this part of the code.
[0020] #include "XD_CMODEL.h" #include".. / lib / vsip.h" #define DEBUG_DIFF long long t_time =0; long long t0=0; long long t1=0; u32 RespType; u32 size; The above is a code example for header file inclusion and general variable definition in the embodiment of this solution. The first two lines are for including header files, the third line defines a macro switch, and the following four lines define general global variables, such as those for timing or statistical quantification of data size, etc. These codes are unified and fixed for each AI model.
[0021] Definition of operator input / output and weight variables In this embodiment, the definition of operator-related variables follows the following rules: Parse the names of each operator according to the onnx structure, and define the variable names of each operator in the inference code according to the operator names. The following is a code example of operator variable definition given in this embodiment: u64 block1_conv1_Conv_out = 0x1080000000 + 134912; u64 block1_conv1_Conv_wb = 0x1080000000 + 245056; u64 block1_conv2_Conv_out = 0x1080000000 + 258144; u64 block1_conv2_Conv_wb = 0x1080000000 + 387616; u64 block1_Add_out = 0x1080000000 + 400448; For a certain model to be deployed, if the name of an operator in its onnx structure is block1_conv1_Conv, after obtaining the operator name, define the weight variable name of this operator by adding the suffix "_wb" to block1_conv1_Conv, and the output variable name by adding the suffix "_out" to block1_conv1_Conv. Since the operator input must be the output of a certain operator, the operator input variable is not defined here. As shown in the above operator variable code example, this solution defines operator-related variables in the form of pointers to achieve the purpose of reasonable memory allocation. "0x1080000000" is the starting memory address, and the offset numbers added later can be calculated according to the output matrix or weight matrix. The difference between the offset number of the next variable and the offset number of the current variable represents the memory size allocated to the current variable. The onnx structure contains the sizes of the output and weight of each operator, so it is relatively convenient to calculate these offsets and generate this part of the code. The sizes here are the original sizes obtained from onnx. If a certain operator needs to pad the output, the size of the output variable during inference is re-determined according to the pad value.
[0022] Loading input file and weight file After the variable definition is completed, load the model data of the model to be deployed, mainly including weight data and model input data. The following is a schematic code for loading the operator weights in the embodiment of this solution: size = get_bin_size(" / ffx0 / models / bins / block1_conv1_Conv_wb.bin"); read_bin(" / ffx0 / models / bins / block1_conv1_conv_wb.bin",block1_conv1_Conv_wb, size); size = get_bin_size(" / ffx0 / models / bins / block1_conv2_conv_wb.bin"); read_bin(" / ffx0 / models / bins / block1_conv2_conv_wb.bin",block1_conv2_Conv_wb, size); Suppose the operator name is parsed from the onnx structure as block1_conv1_Conv in the above schematic code for operator weights. Then the weight data name is defined as adding the suffix "_wb.bin" to block1_conv1_Conv, and the weight variable name is adding the suffix "_wb" to block1_conv1_Conv. Obtain the weight data name and weight variable name according to the operator name. At the same time, generate the first bin function - the get_bin_size function, which is used to simply read the size of the bin file, and the second bin function - the read_bin function, which is used to read the data in the bin file and assign it to the weight variable. In this way, the operator weight loading code is automatically generated through the operator name.
[0023] Inference code generation This embodiment generates inference code according to the onnx structure of the AI model. Common operators included in the onnx model are activation functions such as conv, gemm, relu, pooling functions such as maxpool, and operators such as unsqueeze, flatten, transpose, resize, add, concat. The HC3080 chip provides APIs for some operators, such as activation functions like conv, gemm, relu, and pooling functions like maxpool. For the operators for which the HC3080 chip provides APIs, the APIs provided by the chip are used to generate the inference code for the operators. For the operators for which the HC3080 chip does not provide APIs, this solution provides a set of APIs implemented in C or other feasible languages for generating the inference code for the corresponding operators. The API library used to generate the inference code includes the APIs provided by the AI chip and the set of APIs implemented in C or other feasible languages.
[0024] The process of generating the inference code is to call all the operators in sequence according to the order of the operators in the onnx structure, and the output of the current operator needs to be passed to the next associated operator. In this way, the output of the last operator is the final inference result. The call of all operators will inevitably involve input variables and output variables, and operators with weights will also involve weight variables. These data such as input, output, and weights are stored in the corresponding variables, that is, the memory space allocated in the above "definition of operator input / output and weight variables". The operator API calls the relevant variables according to the association relationship of the operators and stores the operator prediction result in the output variable.
[0025] This solution has made an associated definition of each variable according to the operator name, that is, as Figure 2 shown, for the conv operator, the weight variable name is defined as the operator name conv plus the suffix "_wb". The weight variable name is unique and different for each operator. If there is a weight, it is defined according to the above rules, otherwise the weight variable is not defined. Similarly, the output variable name is defined as the operator name conv plus the suffix "_out". The output variable name is also unique and different for each operator. At the same time, this solution further adds the weight variable name to the weight variable name index table - global_list_wb list according to the order of the operators. For the case where some operators have no weights, 0 is filled in the corresponding position in the list for distinction; similarly, according to the order of the operators, the output variable name is added to the output variable name index table - global_list_out list.
[0026] The input variable name is determined by the output variable name of the associated operator, such as Figure 3As shown, the input variables of different operators may be the same, and an operator may also have multiple input variables. The input variable names of each operator can be determined according to the connection relationship of the operators in the onnx structure. Similarly, add the input variable names to the input variable name index table - the global_list_in list. In this way, the input, output, and weight variable names of the operator can be quickly found according to the operator name or operator index.
[0027] The definition of the general operator API is mostly in the following format: operator_api(input_var,weight_var,output_var,input_shape,output_shape,pad) There may be more parameters, but the following six parameters should be considered first. input_var, weight_var, and output_var are the input variable, weight variable, and output variable respectively. input_shape is the size of the input variable during inference, which determines the data scale involved in the operation. output_shape is the size of the output variable during inference, which must be used when calculating the inference error of this operator mentioned later. At the same time, output_shape is also the input size of the associated operator during inference. The Pad value will change the size of the input or output. Especially on the HC3080, the pad method is different from the conventional method, which is particularly important here.
[0028] Pad generally refers to the padding number of the convolutional layer or pooling layer. For example Figure 4 , when inferring the conv operator with a pad of 1 in the onnx model, the calculation order is to first pad the input to 1*47*34*34, and then send the padded input to the conv-related API for inference. When inferring on the HC3080, if the inference is also carried out in this order, the inference time is relatively long because the pad operation needs to be implemented in C language. Although the conv operator on the HC3080 does not support padding before the operation, it allows padding one or more circles of 0 in the output. When the next operator following this conv operator is exactly a conv operator or a pooling operator (activation functions such as relu can be skipped and do not change the input and output sizes), the padding in the output is exactly for the padding in the next operator, which can achieve the effect of avoiding C language padding. For example Figure 5 , on the HC3080, the pad of the current operator is implemented using the output pad of the previous conv operator, that is Figure 5 the output of the first conv in
[0029] This solution is based on the characteristic of the conv operator on HC3080 that padding is performed after operation. When calling the conv operator, a pad parameter is determined first. This parameter does not describe whether the current operator needs to be padded, but whether the next-level operator or the next few operators associated with it need to be padded and how much padding is required. Therefore, it is necessary to traverse the operators after the current conv operator according to the onnx structure to determine the pad value. The following is a pseudo-code example of the embodiment of this solution for determining the Pad value: Define the function get_pad(inlist, oplist, max_pad); If the first element of inlist contains "conv" Return the padding index (p_INDEX) of inlist If the first element of inlist contains "pool" Return 0 If the next data node (DATA_NEXT) of inlist is empty Return the padding index (p_INDEX) of inlist Traverse all the next data nodes (DATA_NEXT) of inlist Calculate the index j as the current node - 1 (index starts from 0) Recursively call the get_pad function to obtain pad_comp If pad_comp is greater than max_pad Update max_pad to pad_comp Return max_pad The operator list inlist is a list describing the current operator, including the operator name, the pad value (P_INDEX) of the operator in onnx, and the index of the next operator (DATA_NEXT) associated with the current operator. The operator feature list oplist is a list describing all operator features, where each element is an inlist list of a certain operator. max_pad represents the maximum pad value traversed so far. There are three termination conditions for this recursive function, namely, encountering a conv operator, encountering the last operator of the model whose DATA_NEXT is empty, and encountering a pool operator. When encountering a pool operator, return a pad value of 0 because the maxpool operator needs to fill with the minimum value of the input rather than 0, but the conv operator on HC3080 can only fill with 0, so no padding is selected here when encountering a pool operator. After calculating the pad value, add it to the pad value index table - the global_list_real_pad list.
[0030] The pad value of each operator in the model is determined in the above way. For general operators, pad is used to calculate the output variable size during inference. For convolutional and pooling operators, the parameters of the API include this pad value.
[0031] Based on the pad values in the global_list_real_pad list and the operator output information in the onnx model, calculate the output_shape of each operator: H new =H old +2*pad W new =W old +2*pad H new 、W new respectively represent the height and width of the output variable during operator inference; H old 、W old respectively represent the height and width of the operator output variable in onnx.
[0032] Similarly, add them to the output size index table - global_list_real_output list in order. Based on the connection relationship of operators in the onnx structure, similar to obtaining the input variable names, obtain the input_shape of each operator, and similarly add it to the input size index table - global_list_real_input list in order.
[0033] So far, this solution has designed six global lists: global_list_in, global_list_wb, global_list_out, global_list_real_pad, global_list_real_input, and global_list_real_output, which store the input variable names, weight variable names, output variable names, operator pad values, input sizes, and output sizes required for operator calls in the order of operators. According to the above global lists, the parameter information required for operator calls, that is, the key parameters of the API, can be obtained.
[0034] Because operator inference is from front to back, and incorrect calculations of previous operators will affect the inference of subsequent operators. Therefore, this solution adds a function to calculate the inference error of each operator in the deployment code, and quickly confirms whether the inference is correct or which operator has an incorrect inference based on the inference error check of each operator. Figure 6The first four lines are the regular conv operator calls on HC3080, and the last four lines are the code for calculating the inference error. compare_rlt_offset_diff is the corresponding function. The first parameter is the amount of data participating in the statistics, the second parameter indicates how many bytes are used to describe each number, the third parameter indicates the predicted value address, the fourth parameter indicates how many numbers are offset each time, and the fifth parameter indicates the theoretical output value calculated using the onnx model.
[0035] An example of the inference error calculation result in the embodiment of this solution is as follows: compare_rlt start. points 1600 total diff 961 max value of diff 2 cos similiraty : 1.000000 Finally, the compare_rlt_offset_diff function will output the cosine similarity cos_similiraty between the predicted value and the theoretical value, the maximum value max_value_of_diff of the absolute value of the difference between the predicted value and the theoretical value among all points participating in the comparison, and the sum total_diff of the absolute value of the difference between the predicted value and the theoretical value among all points participating in the comparison. Generally, when max_value_of_diff is less than 100 and cos_similiraty is greater than 0.999, the inference error is considered acceptable.
[0036] So far, through the above method, six global lists required for API calls are generated according to the onnx model structure, and combined with the quantization information of the operator, the inference code can be automatically and quickly generated (generally 3 to 9 truncated values, and how to generate the quantization information is not discussed here). For example Figure 6, taking the use of the HC3080 upper convolution API as an example, the information obtained above already includes the parameter information contained in 9 boxes. The first is the operator input and output sizes, which are obtained from the two lists of global_list_real_input and global_list_real_output. According to the order of the current convolution operator, the data of the same order is taken from these two lists. The second is the number of convolution input and output channels, which can be obtained from the onnx file. The third is the pad parameter, which is obtained from the global_list_real_pad list. Similarly, according to the order of the current convolution operator, the data of the same order is taken from this list. The fourth, fifth, and sixth are quantization parameters, which are obtained from the quantization information. The seventh is the input variable name, the eighth is the weight variable name, and the ninth is the output variable name. The seventh, eighth, and ninth parameters are obtained from the three lists of global_list_in, global_list_wb, and global_list_out respectively. The variable names can be taken from these three lists according to the operator name. During inference, the corresponding operator weights are extracted through the weight loading code and stored in the corresponding memory allocation space for inference use. Data, including input data and weight data, is extracted from the memory allocation space of the corresponding variable according to the variable name. At the same time, the inference result (i.e., the output data) of this convolution operator is stored in the memory allocation space of the corresponding variable. The next connected operator with this output as the input extracts data from this memory allocation space. Most of the remaining parameters are fixed and unchanged, and they are called fixed parameters. A few parameters such as the convolution kernel size can be directly obtained from the onnx structure, and they are called operator parameters. The operator parameters are determined according to the specific operator and API. The positions of the parameters are fixed, so the API call code can be automatically and efficiently generated. Combining with the operator inference error calculation logic described above, this part of the code can also be automatically generated. The finally generated model inference code includes each operator API call and operator inference error calculation code.
[0037] File handle release The generation of C code actually writes the above code into a C file using a file handle. Therefore, the file handle must be closed after writing the code. In addition, if the function of automatically generating deployment code is provided to users in the form of a GUI, different model C codes will definitely be repeatedly generated after the program starts. Therefore, after each generation of C code, the contents of the above six global lists are cleared to avoid the operator information of the previous model affecting the generation of the deployment code for the current model. The code examples for file handle release and global list clearing in this embodiment are as follows: c_f.close() global_list_in.clear() global_list_out.clear() global_list_wb.clear() global_list_real_pad.clear() global_list_real_output.clear() global_list_real_input.clear() The inference code implemented above, during the inference process, when it executes to Figure 6 the convolution operator shown, it obtains the weight variable name, input variable name, and output variable name through this convolution API. Then, according to the input variable name and weight variable name, it extracts the input data from the storage space of the corresponding input variable, extracts the weight data from the storage space of the corresponding weight variable, stores the output data in the storage space with the corresponding variable name. After sequentially executing all the operators, the inference ends.
[0038] The specific embodiments described in this article are merely illustrative of the spirit of this solution. Those skilled in the art of this solution can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but they will not deviate from the spirit of this solution or exceed the scope defined by the appended claims.
Claims
1. An automatic inference code generation method for the model deployment process, characterized in that, This method includes: Extracting model information from the intermediate file of the model to be deployed; the model information includes operator names, operator connection relationships, operator pad values, and the original sizes of each variable of the operator; Defining variables for each operator, including defining variable names according to operator names and defining memory allocation for variables according to the original sizes of each variable; The variables of the operator include output variables and weight variables; Loading the model data of the model to be deployed and generating operator weight loading code for each operator; Generating an output variable name index table and a weight variable name index table according to the order of operators in the intermediate file and the definition of each operator variable; Determining the input variable names of each operator according to the connection relationship of the operators, and generating an input variable name index table in the order of the operators; Generating a pad value index table, an input size index table, and an output size index table during actual inference according to the association relationship of the operators in the intermediate file and the model information; For each operator, extracting the corresponding API from the API library, and filling the parameters of the API according to the output variable name index table, the weight variable name index table, the input size index table, the output size index table, the input variable name index table, and the pad value index table to generate the API call of the corresponding operator; Combining the API calls of each operator to obtain the required inference code.
2. The method for automatically generating inference code for the model deployment process according to claim 1, wherein Generating operator weight loading code for each operator in the following way: Defining the weight data name of the corresponding operator according to the operator name; Generating a first bin function for the operator to read the size of the weight data; Generating a second bin function for the operator to read the weight data and assign it to the weight variable.
3. The method for automatically generating inference code for the model deployment process according to claim 2, wherein Defining variable names and weight data names according to operator names specifically includes: For output variables, adding a suffix indicating "output" at the end of the corresponding operator name as the output variable name; For weight variables, adding a suffix indicating "weight" at the end of the corresponding operator name as the weight variable name; For weight data, adding a suffix indicating "weight data" at the end of the corresponding operator name as the weight data name.
4. The method for automatically generating inference code for the model deployment process according to claim 1, characterized in that, Defining the variable names of each operator in the form of pointers to define the memory allocation of the variables while defining the variable names; The generated API calls also include quantization information, fixed parameters, and operator parameters obtained from the intermediate file.
5. The method for automatically generating inference code in the model deployment process according to claim 1, characterized in that The generation of the inference code also includes header file introduction and general variable definition; After generating the inference code, this method further includes closing the file handle used to write the inference code into the C file.
6. The method for automatically generating inference code in the model deployment process according to claim 1, characterized in that Generating a pad value index table, an input size index table, and an output size index table during actual inference according to the association relationship of the operators in the intermediate file and the model information specifically includes: Determining the pad value of the current operator based on the model information, and generating a pad value index table according to the order of operators in the intermediate file; The pad value is used to describe whether the subsequent level operators associated with the current operator need to be filled and the size of the filling number; Obtaining the output size during inference of each operator according to the pad value of each operator in the pad value index table and the original output size of each operator in the intermediate file to obtain the output size index table; Obtain the input size during inference for each operator according to the connection relationships of the operators in the intermediate file to obtain the input size index table described above.
7. The method for automatically generating inference code for the model deployment process according to claim 6, characterized in that, Determine the pad value of the current operator in the following manner: Extract the operator feature list of the intermediate file, where the operator feature list is used to describe all operator features of the currently to-be-deployed model, and each element is an operator list corresponding to the operator; The operator list is used to describe the current operator, including the operator name, the pad value P_INDEX of the operator in the intermediate file, and the index of the subsequent operator associated with the current operator; Traverse the operators after the current operator according to the operator feature list and the operator list, and take the maximum P_INDEX among them as the pad value of the current operator until the convolutional operator, pooling operator, or the last operator of the model is traversed.
8. The method for automatically generating inference code for the model deployment process according to claim 1, wherein This method further includes: Add inference error prediction code including code macro switches to each of the above API calls, which is used to calculate the absolute value of the difference between the inference value and the theoretical value, and close the corresponding API call when the absolute value of the difference is greater than or equal to the preset difference value.
9. The method for automatically generating inference code in the model deployment process according to claim 1, wherein, The intermediate file described above is an ONNX structure model; This method is used to automatically generate inference code for the to-be-deployed model in the HC3080 chip.
10. An inference code automatic generation system for the model deployment process, characterized in that it is used to automatically generate inference code for the to-be-deployed model by the method according to any one of claims 1-9.
Citation Information
Patent Citations
Lightweight microprocessor AI reasoning framework
CN115586958A
Automatic generation device for model reasoning codes
CN118409738A
Automatic quantitative deployment method and system for AI chip supporting operator level operation
CN118605890A
Generative model reasoning service deployment method and device based on AI questions and answers
CN119293154A