Compilation method of binary neural network model, storage medium and computer
By using a custom TVM combination operator, the problem of insufficient TVM compilation support for binary neural network models is solved, enabling rapid and efficient deployment of binary neural network models and reducing the demand for storage and computing resources.
Patent Information
- Application Number
- CN202510479622.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
Existing deep neural network compilers, such as the Tensor Virtual Machine (TVM), lack compilation support for binary neural network models (BNN), resulting in slow and inefficient deployment of binary neural networks.
By customizing one or more dedicated TVM combination operators, including Conv1dBin, BinConv1dBin and BinConv1dGap operators, and integrating multiple standard TVM operators, the binary neural network model is loaded and parsed using the Tensor Virtual Machine to generate an improved Relay format computation graph, and the corresponding functions are called sequentially to generate compiled code.
It enables fast and efficient compilation and deployment of binary neural network models, reducing the consumption of storage space and computing resources.
Smart Images

Figure CN120409578A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of binary neural network models, and in particular to a compilation method, a storage medium, and a computer for a binary neural network model based on a tensor virtual machine.
Background Art
[0002] Currently, in the process of deploying a neural network model to a dedicated chip (Neural Processing Unit, abbreviated as NPU for example), it is usually necessary to convert a pre-trained model into an Open Neural Network Exchange format (abbreviated as ONNX) or a similar data format (such as tflite), and then use an AI (Artificial Intelligence) compiler to adapt to different hardware platforms to improve compatibility. Existing deep neural network compilers such as Tensor Virtual Machine (TVM) lack compilation support for binary neural network models (Binary Neural Network Model, BNN), and are even less able to achieve fast and efficient deployment of binary neural networks.
Summary of the Invention
[0003] One of the objectives of the present invention is to provide a compilation method, a storage medium, and a computer for a binary neural network model based on a tensor virtual machine, which can achieve fast and efficient compilation and deployment of the binary neural network model.
[0004] According to one aspect of the present invention, there is provided a compilation method for a binary neural network model based on a tensor virtual machine, including: using the tensor virtual machine to load and parse a pre-trained binary neural network model to obtain an initial Relay format computation graph corresponding to the binary neural network model, where the initial Relay format computation graph includes a plurality of operations arranged in sequential relay, and where the tensor virtual machine includes a variety of TVM standard operators and one or more custom dedicated TVM combination operators, and each dedicated TVM combination operator includes operations of a variety of TVM standard operators; traversing each operation in the initial Relay format computation graph, and replacing a continuous plurality of operations in the initial Relay format computation graph that match any one dedicated TVM combination operator with the matching dedicated TVM combination operator to obtain an improved Relay format computation graph, where the tensor virtual machine configures a corresponding function for each TVM standard operator and a corresponding function for each dedicated TVM combination operator; the tensor virtual machine sequentially calls the functions corresponding to each TVM standard operator and / or each dedicated TVM combination operator according to the operation order in the improved Relay format computation graph and automatically passes in parameters to obtain the compilation code corresponding to each operator.
[0005] According to another aspect of the present invention, the present invention provides a storage medium storing program instructions, and the program instructions are run to execute the above compilation method.
[0006] According to another aspect of the present invention, the present invention provides a computer, which includes a processor and a memory, the memory stores program instructions, and the processor runs the program instructions to execute the above compilation method.
[0007] Compared with the prior art, the present invention can quickly and efficiently implement the compilation and deployment of a binary neural network model through one or more customized dedicated TVM combined operators, and each dedicated TVM combined operator includes a variety of TVM standard operators.
Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0009] Figure 1 It is a flowchart of a compilation method for a binary neural network model based on a tensor virtual machine in an embodiment of the present invention;
[0010] Figure 2a It is an example of a pre-trained binary neural network model;
[0011] Figure 2b It is to load and parse using the tensor virtual machine Figure 2a The pre-trained binary neural network model of the example to obtain an initial Relay format operation graph;
[0012] Figure 2c It is Figure 2b The improved Relay format operation graph corresponding to the initial Relay format operation graph in;
[0013] Figure 2d It is for compiling Figure 2c The improved Relay format operation graph in to obtain compiled code.
Detailed Embodiments
[0014] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0015] As used herein, "one embodiment" or "an embodiment" refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an isolated or alternative embodiment mutually exclusive with other embodiments.
[0016] The present invention provides a compilation method for a binary neural network model based on a tensor virtual machine, which can quickly and efficiently implement the compilation and deployment of the binary neural network model.
[0017] Figure 1 It is a flowchart of the compilation method for the binary neural network model based on the tensor virtual machine in one embodiment of the present invention. As Figure 1 shown, the conversion method includes the following steps.
[0018] Step 110: Use the tensor virtual machine to load and parse the pre-trained binary neural network model to obtain the initial Relay format operation graph corresponding to the binary neural network model. The initial Relay format operation graph includes multiple operations arranged in sequence in a relay manner. Among them, the tensor virtual machine includes various TVM standard operators and one or more custom dedicated TVM combination operators. All operations in the initial Relay format operation graph are the TVM standard operators, and each dedicated TVM combination operator includes operations of multiple TVM standard operators. The TVM standard operator is each operator supported in the standard Relay format operation graph. Since the TVM standard operator does not involve improvements to the prior art, it will not be elaborated here.
[0019] Step 120: Traverse each operation in the initial Relay format operation graph, and replace consecutive multiple operations in the initial Relay format operation graph that match any one dedicated TVM combination operator with the matching dedicated TVM combination operator to obtain an improved Relay format operation graph. Among them, the tensor virtual machine configures a corresponding function for each TVM standard operator and a corresponding function for each dedicated TVM combination operator.
[0020] Step 130: The tensor virtual machine sequentially calls the functions corresponding to each TVM standard operator and / or each dedicated TVM combination operator according to the operation sequence in the improved Relay format operation graph and automatically passes in parameters to obtain the compilation code corresponding to each operator.
[0021] The compilation method further includes: deploying the compilation code obtained after compiling the binary neural network model to run on a neural network processing unit, where the compilation code corresponding to the compiled dedicated TVM combination operator also runs on the neural network processing unit.
[0022] The dedicated TVM combined operator includes one or more of the following operators: Conv1dBin combined operator, BinConv1dBin combined operator, and BinConv1dGap combined operator.
[0023] Specifically, the Conv1dBin combined operator integrates the Conv operator, MaxPool operator, and Sign operator, and is used to perform convolution, max pooling, and binary quantization on the input feature vector. Among them, the Conv operator, MaxPool operator, and Sign operator are the TVM standard operators.
[0024] The inputs of the Conv1dBin combined operator include:
[0025] An input feature vector with a dimension of L*Ci, where L is the length of the input feature vector and Ci is the number of channels of the input feature vector;
[0026] The weights of the convolution, with a dimension of Co*Ci*K, where Co is the number of output channels, Ci is the number of channels of the input feature vector corresponding to the Ci of the input feature vector, and K is the size of the convolution kernel;
[0027] The bias term of the convolution, with a dimension of Co, where Co is the number of output channels;
[0028] The parameters of convolution, max pooling, and binary quantization, which include: Padding, which determines whether to pad 0 at the beginning and end of the input feature vector; Stride, which determines the convolution stride of the convolution; Pool_size, which determines the stride of the pooling; PackBits, when performing binary quantization on the output, the tensor virtual machine packs PackBits 1-bit data into a single PackBits-bit data, and it is configured as 8, 16, or 32.
[0029] The output of the Conv1dBin combined operator only includes:
[0030] An output feature vector with an output dimension of Lo*[Co / PackBits], where Lo is the length of the output feature vector, and its calculation formula is: Lo = (L + Padding * 2 - (K - 1)) / (stride * Pool_size), and Co / PackBits is the number of channels of the output after packing;
[0031] For the Conv1dBin combined operator, a C function based on a dedicated accelerator is designed. By compiling and running this C function on the dedicated accelerator, the Conv1dBin combined operator can be implemented on the dedicated accelerator, and then the corresponding compiled code of the Conv1dBin combined operator can be obtained.
[0032] Specifically, the BinConv1dBin combined operator fuses the Sign operator, Conv operator, MaxPool operator, and Sign operator, and is used to perform binary quantization, convolution, max pooling, and binary quantization on the input vector. Among them, the Sign operator, Conv operator, MaxPool operator, and Sign operator are the TVM standard operators.
[0033] The inputs of the BinConv1dBin combined operator include:
[0034] An input feature vector with a dimension of L * [Ci / PackBits], where L is the length of the input feature vector, Ci is the number of channels of the input feature vector, and Ci / PackBits is the number of channels of the input feature vector after packing;
[0035] The weights of the convolution, with a dimension of Co * K * [Ci / PackBits], where Co is the number of output channels, Ci is the number of channels of the input feature vector, Ci / PackBits is the number of channels of the input feature vector after packing, and K is the size of the convolution kernel;
[0036] The bias term of the convolution, with a dimension of Co, where Co is the number of output channels;
[0037] The parameters of convolution, pooling, and binary quantization, which include: Padding, which determines whether to pad 0 at the beginning and end of the input feature vector; stride, which determines the convolution stride of the convolution; Pool_size, which determines the stride of the pooling; PackBits, when performing binary quantization on the output, the tensor virtual machine packs PackBits 1-bit data into a single PackBits-bit data, and can be configured as 8, 16, or 32.
[0038] In the Conv1dBin combined operator, in order to reduce the storage space of the binary convolution weights, the original convolution kernel weights are packed. Specifically: the original convolution kernel weights are Co * K * Ci, and along the Ci direction, every PackBits weight data is packed into a PackBits-bit weight data. That is, the dimension of the new convolution weights is Co * K * [Ci / PackBits].
[0039] The output of the BinConv1dBin combined operator only includes:
[0040] An output feature vector with an output dimension of Lo * [Co / PackBits],
[0041] where Lo is the length of the output feature vector, and its calculation formula is:
[0042] Lo = (L + Padding * 2 - (K - 1)) / (stride * Pool_size),
[0043] Co / PackBits is the number of channels of the output after packing.
[0044] For the BinConv1dBin combined operator, a C function based on a dedicated accelerator is designed. By compiling and running this C function on the dedicated accelerator, the BinConv1dBin combined operator can be implemented on the dedicated accelerator, and the corresponding compiled code of the BinConv1dBin combined operator can be obtained.
[0045] Specifically, the BinConv1dGap combined operator integrates the Sign operator, the Conv operator, the Sum operator, and the Reshape operator, and is used for binary quantization, convolution, element summation, and dimension conversion of the input vector. Among them, the Sign operator, the Conv operator, the Sum operator, and the Reshape operator are the TVM standard operators.
[0046] The inputs of the BinConv1dGap combined operator include:
[0047] An input feature vector with a dimension of L * [Ci / PackBits], where L is the length of the input feature vector, Ci is the number of channels of the input feature vector, and Ci / PackBits is the number of channels of the input feature vector after packing;
[0048] The weights of the convolution with a dimension of Co * K * [Ci / PackBits], where Co is the number of output channels, Ci / PackBits is the number of channels of the input feature vector after packing, and K is the size of the convolution kernel,
[0049] The bias term of the convolution with a dimension of Co, where Co is the number of output channels;
[0050] The parameters of the convolution include: Padding, which determines whether to pad 0 at the beginning and end of the input feature vector; stride, which determines the convolution stride of the convolution.
[0051] In the BinConv1dGap combined operator, in order to reduce the storage space of the binary convolution weights, the original convolution kernel weights are packed. Specifically: the original convolution kernel weights are Co * K * Ci, and along the Ci direction, every PackBits weight data is packed into a PackBits-bit weight data.
[0052] The output of the BinConv1dGap combined operator only includes:
[0053] Output a feature vector with an output dimension of Co, where Co is the number of output channels.
[0054] For the BinConv1dGap combined operator, a C function based on a dedicated accelerator is designed. By compiling and running this C function on the dedicated accelerator, the BinConv1dGap combined operator can be implemented on the dedicated accelerator, and the corresponding compilation code of the BinConv1dGap combined operator can be obtained.
[0055] In the present invention, the data packing method that may be adopted for the input feature vector, output feature vector, and / or convolution kernel weight in the combined operator. Specifically, the data packing method is as follows:
[0056] Compare the original data with 0. If it is > 0, set its value to 1; otherwise, set it to 0.
[0057] Then, according to the set PackBits, every PackBits fixed-point numbers are packed into 1 fixed-point number with PackBits bits. Taking PackBits = 8 as an example:
[0058] [-5, 100, -35, 50, 88, -60, 33, 0] ->
[0059] [0, 1, 0, 1, 1, 0, 1, 0] ->
[0060] 8’b01011010 (corresponding to the decimal value of 90).
[0061] Figure 2b To use the tensor virtual machine to load and parse Figure 2a the pre-trained binary neural network model in the example to obtain an initial Relay format operation graph; Figure 2c Figure 2c For Figure 2b the improved Relay format operation graph corresponding to the initial Relay format operation graph in ; The improved Relay format operation graph corresponding to the initial Relay format operation graph in Figure 2d For the compilation Figure 2c The compilation code obtained by compiling the improved Relay format operation graph in
[0062] In this way, the present invention realizes the compilation of multiple TVM standard operators quickly and efficiently through one or more custom dedicated TVM combined operators, and each dedicated TVM combined operator includes multiple TVM standard operators. Moreover, the compiled compilation code can occupy less storage space and computing resources during deployment and operation, thereby realizing the quick and efficient compilation and deployment of the binary neural network model.
[0063] According to another aspect of the present invention, the present invention provides a storage medium storing program instructions, and when the program instructions are executed, they are run to execute the compilation method described above. For simplicity, the specific content of the compilation method of the binary neural network model based on the tensor virtual machine is not repeated here.
[0064] According to another aspect of the present invention, the present invention provides a computer, which includes a processor and a memory. Program instructions are stored in the memory, and the processor runs the program instructions to execute the conversion method described above or the deployment method described above. For simplicity, the specific content of the compilation method of the binary neural network model based on the tensor virtual machine is not repeated here.
[0065] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0066] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications and variations to the above embodiments within the scope of the present invention.
Claims
1. A compilation method for a binary neural network model based on a tensor virtual machine, characterized in that It includes: Loading and parsing a pre-trained binary neural network model by using the tensor virtual machine to obtain an initial Relay format operation graph corresponding to the binary neural network model, where the initial Relay format operation graph includes a plurality of operations arranged in sequential relay, and the tensor virtual machine includes a variety of TVM standard operators and one or more dedicated TVM combined operators defined by oneself, and each dedicated TVM combined operator includes operations of a variety of TVM standard operators; Traversing each operation in the initial Relay format operation graph, and replacing a continuous plurality of operations in the initial Relay format operation graph that match any one dedicated TVM combined operator with the matching dedicated TVM combined operator to obtain an improved Relay format operation graph, where the tensor virtual machine configures a corresponding function for each TVM standard operator and a corresponding function for each dedicated TVM combined operator; and The tensor virtual machine sequentially calls the functions corresponding to each TVM standard operator and / or each dedicated TVM combined operator according to the operation sequence in the improved Relay format operation graph and automatically passes in parameters to obtain the compilation code corresponding to each operator.
2. The compilation method according to claim 1, wherein It further includes: Deploying the compilation code obtained after compiling the binary neural network model to run on a neural network processing unit, where the compilation code corresponding to the compiled dedicated TVM combined operator also runs on the neural network processing unit.
3. The compilation method according to claim 1, wherein The dedicated TVM combined operator includes one or more of the following operators: Conv1dBin combined operator; BinConv1dBin combined operator; BinConv1dGap combined operator.
4. The compilation method according to claim 3, wherein The Conv1dBin combined operator integrates a Conv operator, a MaxPool operator, and a Sign operator, and is used for performing convolution, max pooling, and binary quantization on an input feature vector. The inputs of the Conv1dBin combined operator include: An input feature vector, whose dimension is L*Ci, where L is the length of the input feature vector and Ci is the number of channels of the input feature vector; The weight of convolution, whose dimension is Co*Ci*K, where Co is the number of output channels, Ci is the number of channels of the input feature vector corresponding to the Ci of the input feature vector, and K is the size of the convolution kernel; The bias term of convolution, whose dimension is Co, and Co is the number of output channels; The parameters of convolution, max pooling, and binary quantization, which include: Padding, which determines whether to pad 0 at the beginning and end of the input feature vector; Stride, which determines the convolution stride of convolution; Pool_size, which determines the stride of pooling; PackBits, when performing binary quantization on the output, the tensor virtual machine packs PackBits 1-bit data into a single data of PackBits bits, and it is configured as 8, 16, or 32; The output of the Conv1dBin combined operator only includes: Output feature vector, whose output dimension is Lo * [Co / PackBits], where Lo is the length of the output feature vector, and its calculation formula is: Lo = (L + Padding * 2 - (K - 1)) / (stride * Pool_size), and Co / PackBits is the number of channels of the output after packing; For the Conv1dBin combined operator, a C function based on a dedicated accelerator is designed. By compiling and running this C function on the dedicated accelerator, the Conv1dBin combined operator can be implemented on the dedicated accelerator, and thus the corresponding compilation code of the Conv1dBin combined operator can be obtained.
5. The compilation method according to claim 3, wherein The BinConv1dBin combined operator combines the Sign operator, the Conv operator, the MaxPool operator and the Sign operator, and is used to perform binary quantization, convolution, max pooling and binary quantization on the input vector; The input of the BinConv1dBin combined operator includes: Input feature vector, whose dimension is L * [Ci / PackBits], where L is the length of the input feature vector, Ci is the number of channels of the input feature vector, and Ci / PackBits is the number of channels of the input feature vector after packing; The weight of the convolution, whose dimension is Co * K * [Ci / PackBits], where Co is the number of output channels, Ci is the number of channels of the input feature vector, Ci / PackBits is the number of channels of the input feature vector after packing, and K is the size of the convolution kernel; The bias term of the convolution, whose dimension is Co, and Co is the number of output channels; The parameters of convolution, pooling and binary quantization, which include: Padding, which determines whether to pad 0 at the beginning and end of the input feature vector; stride, which determines the convolution stride of the convolution; Pool_size, which determines the stride of the pooling; PackBits, when performing binary quantization on the output, the tensor virtual machine packs PackBits 1-bit data into a single PackBits-bit data, and can be configured as 8, 16 or 32; The output of the BinConv1dBin combined operator only includes: Output feature vector, whose output dimension is Lo * [Co / PackBits], where Lo is the length of the output feature vector, and its calculation formula is: Lo = (L + Padding * 2 - (K - 1)) / (stride * Pool_size), Co / PackBits is the number of channels of the output after packing; For the BinConv1dBin combined operator, a C function based on a dedicated accelerator is designed. By compiling and running this C function on the dedicated accelerator, the BinConv1dBin combined operator can be implemented on the dedicated accelerator, and thus the corresponding compilation code of the BinConv1dBin combined operator can be obtained.
6. The compilation method according to claim 3, wherein The BinConv1dGap combined operator integrates the Sign operator, Conv operator, Sum operator, and Reshape operator, and is used for binary quantization, convolution, element summation, and dimension conversion of the input vector; The inputs of the BinConv1dGap combined operator include: An input feature vector with a dimension of L * [Ci / PackBits], where L is the length of the input feature vector, Ci is the number of channels of the input feature vector, and Ci / PackBits is the number of channels of the input feature vector after packing; The weights of the convolution, with a dimension of Co * K * [Ci / PackBits], where Co is the number of output channels, Ci / PackBits is the number of channels of the input feature vector after packing, and K is the size of the convolution kernel, The bias term of the convolution, with a dimension of Co, where Co is the number of output channels; The parameters of the convolution, which include: Padding, which determines whether to pad zeros at the beginning and end of the input feature vector; Stride, which determines the convolution stride of the convolution; The output of the BinConv1dGap combined operator only includes: An output feature vector with an output dimension of Co, where Co is the number of output channels, For the BinConv1dGap combined operator, a C function based on a dedicated accelerator is designed. By compiling and running this C function on the dedicated accelerator, the BinConv1dGap combined operator can be implemented on the dedicated accelerator, and the corresponding compilation code of the BinConv1dGap combined operator can be obtained.
7. A storage medium, characterized in that, It stores program instructions, and the program instructions are run to execute the compilation method according to any one of claims 1-6.
8. A computer, characterized in that, It includes a processor and a memory. The memory stores program instructions, and the processor runs the program instructions to execute the compilation method according to any one of claims 1-6.