Convolutional neural network compiling method based on CPU / NPU cooperative computing

In the CPU/NPU collaborative computing environment, the fusion of the ONNX-MLIR toolchain and the NPU compiler is used to realize the quantization and compression of the CNN model, and through technologies such as task scheduling, operator packaging and memory management, the computing performance of CNN in heterogeneous architecture is optimized, solving the problem that CNN cannot operate effectively in the existing technology.

CN119938057AActive Publication Date: 2025-05-06SHANGHAI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510014346.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

When the existing technology uses CPU/NPU collaborative computing, the lack of mature architecture and toolchain support has led to the inability of CNN to operate effectively on heterogeneous architectures (RISC-V CPU/NPU), and there is a computing performance bottleneck.

Method used

A convolutional neural network compilation method based on CPU/NPU collaborative computing is proposed. Through the fusion of the ONNX-MLIR toolchain and the NPU compiler, the quantization and compression of the CNN model is realized, and the task list allocated to the CPU and NPU through the technologies such as task scheduling, operator packaging and memory management is generated to optimize the computing performance.

Benefits of technology

Through the deep integration and optimization of CPU and NPU, the computing performance of CNN in heterogeneous architecture has been significantly improved, solving the problem that CNN cannot operate effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938057A_ABST
    Figure CN119938057A_ABST
Patent Text Reader

Abstract

The invention discloses a convolutional neural network compiling method based on CPU / NPU cooperative computing, and the method comprises the steps: carrying out the model quantification and compression of various CNN models under different network frameworks, and exporting the models in an ONNX file format; after an ONNX file compiling process is completed through fusion of an ONNX-MMIR tool chain and an NPU compiler and task lists distributed to a CPU and an NPU are generated respectively, at the CPU end, the ONNX-MMIR generates an effective RISC-V instruction for a CPU task, efficient scheduling is supported, and at the NPU end, the customized NPU compiler completes matching of the NPU task on algorithms and hardware and efficiently maps operators to an NPU platform. According to the method, the problem that the CNN cannot effectively run on the heterogeneous architecture (RISC-V CPU / NPU) is solved through the performance function and optimization of the core stage, and the calculation performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of information processing, specifically a convolutional neural network compilation method based on CPU / NPU collaborative computing. Background Art

[0002] In recent years, with the rapid development of convolutional neural networks (CNN), neural network processing units (NPUs) have been widely used due to their high energy efficiency and low latency. However, faced with increasingly complex algorithm models, NPUs have bottlenecks in processing performance, while traditional central processing units (CPUs), despite their versatility, have low computing efficiency. Combining the flexibility of CPUs with the efficiency of NPUs provides an effective way to solve these problems. Existing research on CPU / NPU collaborative computing is relatively limited, lacking mature architecture and tool chain support. Summary of the invention

[0003] In view of the above-mentioned deficiencies in the prior art, the present invention proposes a convolutional neural network compilation method based on CPU / NPU collaborative computing, which solves the problem that CNN cannot run effectively on heterogeneous architecture (RISC-V CPU / NPU) by functional functions and optimizing the core stage, and achieves a significant improvement in computing performance.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a convolutional neural network compilation method based on CPU / NPU collaborative computing, comprising:

[0006] Step 1: Quantize and compress various CNN models under different network frameworks and export them in ONNX file format;

[0007] Step 2: Complete the ONNX file compilation process through the fusion of the ONNX-MLIR toolchain and the NPU compiler and generate task lists assigned to the CPU and NPU respectively, including:

[0008] 2.1 In the ONNX dialect stage, operations are allocated between the CPU and NPU according to the specific needs of users through task scheduling functions and three different modes;

[0009] 2.2 In the MLIR dialect stage, the operator packaging technology is used to encapsulate the CNN operators that belong to the CPU tasks divided by the task scheduling function, and multiple operators are combined into more efficient execution units to significantly reduce the overhead of operation switching between the CPU and NPU;

[0010] 2.3 In the LLVM dialect stage, the memory management of heterogeneous platforms is unified through memory management functions and the memory space is compressed through memory optimization to ensure the correctness of data interaction.

[0011] Step 3: Realize the deep integration of CPU and NPU through RISC-V instruction extension. Specifically, on the CPU side, ONNX-MLIR generates valid RISC-V instructions for CPU tasks to support efficient scheduling. On the NPU side, the customized NPU compiler completes the matching of NPU tasks in algorithms and hardware and efficiently maps operators to the NPU platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flow chart of the present invention;

[0013] Figure 2 Packing flowchart for operators;

[0014] Figure: (a) Input model; (b) Downgrading to ONNX dialect; (c) Search and operator encapsulation; (d) Mapping to hardware

[0015] Figure 3 Common NPU memory optimization methods;

[0016] In the figure: (a) Resnet50 (first stage BTNK2, the original input feature map size is 224×224×3); (b) read-write conflict in the dual-end memory management mechanism; (c) use-after-release memory management mechanism. DETAILED DESCRIPTION

[0017] like Figure 1 As shown, this embodiment relates to a convolutional neural network compilation method based on CPU / NPU collaborative computing, including:

[0018] Step 1: Quantize and compress various CNN models under different network frameworks and export them in ONNX file format;

[0019] Step 2: Complete the ONNX file compilation process through the fusion of the ONNX-MLIR toolchain and the NPU compiler and generate task lists assigned to the CPU and NPU respectively, including:

[0020] 2.1 In the ONNX dialect stage, operations are allocated between the CPU and NPU according to the specific needs of users through task scheduling functions and three different modes.

[0021] The three different modes include: Full CPU mode, User defined mode and Default mode

[0022] The full CPU mode means that all tasks are forced to be executed on the CPU, regardless of whether they are compatible with the NPU. This mode is mainly used to verify the correctness of calculations, and to ensure the accuracy and consistency of tasks by comparing the calculation results of the CPU and NPU. For scenarios that require high-confidence verification, this mode can provide reliable support.

[0023] The user-defined mode means that advanced users are provided with the ability to manually control task allocation. Users can specify task allocation through a binary string, where each bit represents the task allocation position of an operator: "0" indicates allocation to the CPU, and "1" indicates allocation to the NPU. This mode allows users to flexibly adjust the allocation strategy according to specific needs and hardware characteristics, and is used to evaluate the acceleration effect of the CPU and NPU under different configurations, providing the possibility of in-depth analysis for performance optimization.

[0024] The default mode means that tasks are automatically assigned according to the capabilities of the NPU. The scheduler checks whether each operation meets the constraints of the NPU. If so, it is assigned to the NPU, otherwise it is assigned to the CPU. This mode makes full use of the computing power of the NPU to optimize system performance. In addition, users can adjust the range of tasks supported by the NPU by modifying the "NPU implementable operators" list to analyze the performance of specific operations on the NPU. This mode is particularly suitable for users who want to balance performance and flexibility.

[0025] 2.2 In the MLIR dialect stage, the operator packaging technology is used to encapsulate the CNN operators that belong to the CPU tasks divided by the task scheduling function, and multiple operators are combined into more efficient execution units to significantly reduce the overhead of operation switching between the CPU and NPU, including:

[0026] ① Add custom operators: define new operators in ONNX dialect, including name, input, output and attributes, and ensure that they can be recognized and parsed by the compiler; reduce high-level operators to MLIR dialect and optimize their representation to adapt to hardware requirements; convert operators to LLVM dialect to make them compatible with the LLVM backend and generate efficient machine code. The packaged operators do not need to contain detailed operator or parameter information, but the information of input and output variables must be retained to ensure integrity during dialect conversion.

[0027] ②Introduce three custom operators in ONNX dialect to handle operations with different numbers of input variables, and complete operator optimization through a customized CNN compiler.

[0028] Take the MNIST network as an example. Figure 3As shown in (b), the input computational graph consists of a series of ONNX operators (such as Maxpool, Reshape, Gemm, ReLU, and Softmax). During the downgrading process of ONNX-MLIR, these operators are mapped to ONNX dialect and represented as function calls, preserving the structure of the computational graph. In further optimization, operators such as Gemm and ReLU are packaged as NPU functions, while operators such as Maxpool and Reshape are assigned to the CPU. These operators are finally executed in sequence between the CPU and NPU, and efficient data transmission is achieved through the system bus.

[0029] The operator optimization includes: for the CPU, adjacent operators will be merged into one function to reduce function calls and data transmission overhead; for the NPU, operator packing focuses on maximizing throughput, and common combinations include: convolution + activation function or pooling operation; convolution + addition + activation or pooling operation, which is used to optimize the execution efficiency of the residual network; matrix multiplication + activation or data transformation.

[0030] 2.3 In the LLVM dialect stage, the memory management of heterogeneous platforms is unified through memory management functions and the memory space is compressed through memory optimization to ensure the correctness of data interaction.

[0031] The memory optimization includes: a dual-end (DE) mechanism for configuring a self-checking function and a use-after-release (RU) mechanism for simplifying compilation requirements on certain networks.

[0032] like Figure 3 As shown in (c), taking the 224×224×3 original input feature map of Resnet50 as an example, when NPU2 finishes calculating, the output feature map of NPU1 is released. However, when calculating NPU3, a read-write conflict occurs due to insufficient space, resulting in an error in the calculation result of NPU3.

[0033] The dual-end (DE) mechanism refers to: traversing the total storage space of input feature map data and output feature map data of all layers, and setting the maximum value as the current feature map storage space. When a read-write conflict occurs, the size of the initially allocated memory space is updated through self-checking.

[0034] Step 3: Realize the deep integration of CPU and NPU through RISC-V instruction extension. Specifically: On the CPU side, ONNX-MLIR generates effective RISC-V instructions for CPU tasks, supporting efficient scheduling. On the NPU side, the customized NPU The compiler completes the matching of NPU tasks in algorithm and hardware and efficiently maps operators to the NPU platform. In addition, RISC-V instruction extension, as a key link in the hardware layer, provides core support for the unified operation of heterogeneous instruction sets.

[0035] After specific practical experiments, the present invention is deployed in a CPU / NPU heterogeneous computing system in a Python environment, and the heterogeneous instruction sets of CPU and NPU are integrated through RISC-V instruction extension, combining the versatility of CPU with the high efficiency of NPU, supporting cross-platform collaborative computing of complex CNN operators, improving the overall performance and resource utilization of heterogeneous architecture, solving the problem that CNN cannot run effectively on heterogeneous architecture (RISC-V CPU / NPU), and achieving a significant improvement in computing performance.

[0036] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.

Claims

1. A convolutional neural network compilation method based on CPU / NPU collaborative computing, characterized in that: include: Step 1: Quantize and compress various CNN models under different network frameworks and export them in ONNX file format; Step 2: Complete the ONNX file compilation process through the fusion of ONNX-MLIR toolchain and NPU compiler and generate task lists assigned to CPU and NPU respectively, including: 2.1 In the ONNX dialect stage, operations are allocated between the CPU and NPU according to the specific needs of users through task scheduling functions and three different modes; 2.2 In the MLIR dialect stage, the operator packaging technology is used to encapsulate the CNN operators that belong to the CPU tasks divided by the task scheduling function, and multiple operators are combined into more efficient execution units to significantly reduce the overhead of operation switching between the CPU and NPU; 2.3 In the LLVM dialect stage, the memory management of heterogeneous platforms is unified through memory management functions and the memory space is compressed through memory optimization to ensure the correctness of data interaction. Step 3: Realize the deep integration of CPU and NPU through RISC-V instruction extension. Specifically, on the CPU side, ONNX-MLIR generates valid RISC-V instructions for CPU tasks to support efficient scheduling. On the NPU side, the customized NPU compiler completes the matching of NPU tasks in algorithms and hardware and efficiently maps operators to the NPU platform.

2. The convolutional neural network compilation method based on CPU / NPU collaborative computing according to claim 1 is characterized in that: The three different modes include: full CPU mode, user defined mode and default mode; The full CPU mode means: forcing all tasks to be executed on the CPU; The user-defined mode means that: advanced users are provided with the ability to manually control task allocation; The default mode means that tasks are automatically allocated according to the capabilities of the NPU, that is, the scheduler checks whether each operation meets the constraints of the NPU, and if so, allocates it to the NPU, otherwise allocates it to the CPU.

3. The convolutional neural network compilation method based on CPU / NPU collaborative computing according to claim 1 is characterized in that: The step 2.2 specifically includes: i) Add custom operators: define new operators in ONNX dialect, including name, input, output and attributes, and ensure that they can be recognized and parsed by the compiler; reduce high-level operators to MLIR dialect and optimize their representation to adapt to hardware requirements; convert operators to LLVM dialect to make them compatible with LLVM backend and generate efficient machine code. The packaged operators do not need to contain detailed operator or parameter information, but the information of input and output variables must be retained to ensure integrity during dialect conversion; ii) Three custom operators are introduced in ONNX dialect to handle operations with different numbers of input variables, and the operator optimization is completed through a customized CNN compiler.

4. The convolutional neural network compilation method based on CPU / NPU collaborative computing according to claim 1 is characterized in that: The operator optimization includes: for the CPU, adjacent operators will be merged into one function to reduce function calls and data transmission overhead; for the NPU, operator packing focuses on maximizing throughput, and common combinations include: convolution + activation function or pooling operation; convolution + addition + activation or pooling operation, which is used to optimize the execution efficiency of the residual network; matrix multiplication + activation or data transformation.

5. The convolutional neural network compilation method based on CPU / NPU collaborative computing according to claim 1 is characterized in that: The memory optimizations include: a dual-end mechanism for configuring self-checking functions and a use-after-free mechanism for simplifying compilation requirements on certain networks.

6. The convolutional neural network compilation method based on CPU / NPU collaborative computing according to claim 5 is characterized in that: The dual-end mechanism refers to: traversing the total storage space of input feature map data and output feature map data of all layers, and setting the maximum value as the current feature map storage space. When a read-write conflict occurs, the size of the initially allocated memory space is updated through self-checking.

Citation Information

Patent Citations

  • Method and system for dynamically scheduling multi-core fusion computing processor

    CN114371933A

  • Task compiling method and device and compiler

    CN114911610A

  • Computing power flexible combination method and system based on embedded platform

    CN116737397A

  • Inference task scheduling method and device and computer equipment

    CN116841706A

  • Control of scheduling dependencies by a neural network compiler

    US20190391796A1