Method and system for generating deep learning model for efficient operation of embedded neural processing unit

The method addresses compatibility and resource optimization challenges for deep learning models on embedded NPUs by converting networks to ONNX graphs, generating CFG IR, and optimizing NPU resource allocation, resulting in efficient and compatible deep learning operations on embedded devices.

WO2025095158A1PCT designated stage expired Publication Date: 2025-05-08KOREA ELECTRONICS TECH INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/017073
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2023-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing deep learning models face challenges in achieving high compatibility and optimized operations on embedded devices with built-in Neural Processing Units (NPUs), due to limited resources and framework compatibility issues.

Method used

A method is developed to create a deep learning model by converting the network to an ONNX graph, sorting it into a directional non-cycle graph, extracting layer parameters, and generating CFG IR for optimized NPU resource allocation and binary code generation, ensuring high compatibility and efficient operations.

Benefits of technology

The solution enhances compatibility between deep learning networks and embedded NPUs, optimizing resource allocation and balancing NPU performance and energy consumption, thereby enabling efficient learning and reasoning operations on embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023017073_08052025_PF_FP_ABST
    Figure KR2023017073_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and a system for generating a deep learning model for an efficient operation of an embedded NPU. The method for generating a deep learning network model for an embedded device, according to an embodiment of the present invention, comprises: converting a deep learning network into an ONNX graph; arranging the converted ONNX graph as a directed acyclic graph; extracting a parameter for each layer of the arranged ONNX graph; generating a CFG IR on the basis of the extracted parameter; and on the basis of the generated CFG IR, generating firmware having a binary code capable of controlling training and inference operations through an NPU and data regarding hardware-specific resource allocation of an NPU embedded device. Accordingly, when an embedded device having an NPU embedded therein performs various deep learning network training and inference operations, it is possible to secure high compatibility, and maintain an appropriate balance between performance and energy consumption of the NPU by performing an optimized operation.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for creating a deep learning model for efficient operation of an embedded neural processing unit.

[0001] The present invention relates to a technology for generating a deep learning model, and more particularly, to a method for enabling an embedded device with a built-in NPU (Neural Processing Unit) to secure high compatibility and perform optimized operations when performing various deep learning network learning and inference operations.

[0002] As artificial intelligence networks utilizing deep learning develop, network training and inference are processed through various types of platforms, such as data centers, graphics cards, and AI accelerator chips.

[0003] In particular, recent AI networks are often used in mobile embedded environments equipped with built-in AI accelerators. While other platforms tend to focus on improving computational performance due to ample available hardware resources, embedded environments face challenges in improving performance and achieving learning due to limited resources.

[0004] Additionally, when artificial intelligence networks are developed through various deep learning frameworks such as TensorFlow, Caffe, and PyTorch, there is a problem of poor compatibility when processing various types of networks in mobile embedded environments if the same operators are not corresponding in different frameworks.

[0005] The present invention has been made to solve the above problems, and the purpose of the present invention is to provide a method for generating a deep learning model through data reuse and optimization of NPU resource allocation in a compiler that increases compatibility between various types of designed deep learning networks and embedded NPUs, as a method for ensuring high compatibility and optimized operation performance when an embedded device with a built-in NPU performs various deep learning network learning and inference operations.

[0006] According to one embodiment of the present invention for achieving the above object, a method for generating a deep learning network model for an embedded device includes the steps of: converting a deep learning network into an ONNX graph; aligning the converted ONNX graph into a directed acyclic graph; extracting parameters for each layer of the aligned ONNX graph; generating a CFG IR (Intermediate Representation) based on the extracted parameters; and generating a binary code capable of controlling learning and inference operations through an NPU based on the generated CFG IR and firmware having data on hardware-specific resource allocation of an NPU embedded device.

[0007] CFG IR may include parameters for each layer required for convolution operations, batch normalization operations, activation and pooling function operations, and parameters for computing hardware allocation resources.

[0008] The method for generating a deep learning network model for an embedded device according to the present invention may further include a step of transferring a deep learning network and firmware generated in the generation step to the embedded device.

[0009] Hardware-specific optimized resource allocation of an NPU embedded device may include a process of grouping input data, dividing it into multiple partial data, and then distributing the grouped data to an NPU having a systolic array structure.

[0010] Optimized resource allocation for each hardware of an NPU embedded device may be to minimize the number of times weights are brought into the NPU's internal buffer for data reuse.

[0011] Optimized resource allocation for each hardware of an NPU embedded device may be to specify the number of PEs from the form of segmented input data.

[0012] The hardware-optimized resource allocation of the NPU embedded device may enable the NPU of the NPU embedded device to access memory addresses for data reuse when performing the learning process.

[0013] The data to be reused may be data calculated during the forward propagation process.

[0014] Optimized resource allocation for each hardware of an NPU embedded device may involve allocating memory to store data generated in the backpropagation process performed through data recycling.

[0015] According to another aspect of the present invention, a system for generating a deep learning network model for an embedded device is provided, comprising: a processor for converting a deep learning network into an ONNX graph, aligning the converted ONNX graph into a directed acyclic graph, extracting parameters for each layer of the aligned ONNX graph, generating a CFG IR (Intermediate Representation) based on the extracted parameters, and generating a binary code capable of controlling learning and inference operations through an NPU based on the generated CFG IR and firmware having data on hardware-specific resource allocation of an NPU embedded device; and a storage unit for providing storage space required by the processor.

[0016] According to another aspect of the present invention, a method for generating a deep learning network model for an embedded device is provided, comprising: a step of generating a deep learning network and converting the deep learning network into an ONNX graph; a step of generating a firmware having binary code capable of controlling learning and inference operations through an NPU and data on hardware-specific resource allocation of an NPU embedded device from the converted ONNX graph; and a step of transferring the deep learning network and the firmware generated in the generation step to the embedded device.

[0017] According to another aspect of the present invention, a deep learning network model generation system for an embedded device is provided, comprising: a processor for generating a deep learning network, converting the deep learning network into an ONNX graph, and generating a binary code capable of controlling learning and inference operations through an NPU from the converted ONNX graph and firmware having data on hardware-specific resource allocation of an NPU embedded device; and a communication unit for transmitting the deep learning network and the firmware generated in the generation step to the embedded device.

[0018] As described above, according to embodiments of the present invention, by generating a deep learning model through data reuse and optimization of NPU resource allocation in a compiler that enhances compatibility between deep learning networks designed in various types and embedded NPUs, it is possible to secure high compatibility and maintain an appropriate balance between NPU performance and energy consumption through optimized computation when an embedded device with a built-in NPU performs various deep learning network learning and inference operations.

[0019] Figure 1. NPU Global Market Size Forecast Research Data (MarketsandMarkets, 2023)

[0020] Figure 2. Structure of an embedded device equipped with an NPU

[0021] Figure 3. Interconnection structure between a host (PC) equipped with an NPU compiler and an embedded device.

[0022] Figure 4. Data reuse and additional allocation of backward data addresses during the learning process.

[0023] Figure 5. Hardware configuration of a deep learning model generation system according to another embodiment of the present invention.

[0024] Hereinafter, the present invention will be described in more detail with reference to the drawings.

[0025] Advanced deep learning networks, with their complex structures and massive sizes, demand high hardware performance to handle them. However, with the expected increase in operating costs for existing data center server-based AI services to process these deep learning networks, efforts are being made to address this issue through low-power, low-cost NPUs specialized for deep learning network computation.

[0026] Figure 1 presents data on the size and growth trends of the global NPU market. It is projected to grow at an annual rate of over 28.1% through 2028, reaching a market size of approximately $64.5 billion. As the number of users utilizing NPU-based deep learning network services increases, mobile embedded NPU systems, which can perform not only model inference but also training by embedding NPUs in the mobile devices closest to the user, are also attracting attention.

[0027] Unlike traditional server-based deep learning network processing, embedded NPUs face limitations due to limited hardware resources. Therefore, compatibility and scalability between hardware (available resources within the mobile device) and software (optimized network and firmware), along with efficient computational processing, must be considered for compilers that generate executable code to enable optimized NPU operation.

[0028] In an embodiment of the present invention, a method for generating a deep learning model using a compiler that can maximize large-scale learning and inference processing performance in a mobile embedded NPU system is proposed.

[0029] This is a firmware generation and hardware resource allocation technology to ensure high compatibility of embedded NPUs and perform optimized operations when embedded devices with built-in NPUs perform various deep learning network training and inference operations.

[0030] Figure 2 illustrates the structure of a mobile embedded device to which an embodiment of the present invention can be applied. The embedded device comprises a CPU (130) as a main processing unit, an NPU (140) for processing deep learning network operations, a DRAM (120) for storing input / output data required for the operation process, and an external input / output interface (110).

[0031] The NPU (140) receives data for processing deep learning network operations from the DRAM (110), stores it in the On-chip Buffer (141), and performs the operations in the PE (Processing Engine, 142) optimized for the operations, and the results are written back to the DRAM (110).

[0032] However, the on-chip buffer (141) and PE (142) within the NPU (140), including the DRAM (120), are designed within limited resources, requiring a process to minimize wasted resources and optimize the computational flow. To address this, a compiler is required that generates optimized executable code for the pre-designed deep learning network to operate at maximum performance in a mobile embedded environment.

[0033] Figure 3 is a diagram showing the overall structure in which an NPU compiler and an embedded NPU device are combined.

[0034] The compiler (200) first converts a deep learning network designed for learning and inference on an embedded device into the ONNX (Open Neural Network Exchange) format. When using deep learning networks designed with different frameworks, the same computational layers may be expressed in different ways, potentially leading to poor compatibility. This problem is resolved by converting to the ONNX format.

[0035] The following compiler (200) sorts the converted ONNX graph into a directed acyclic graph, extracts parameters for each layer of the sorted ONNX graph, and generates a CFG IR (CinFiGuraition Intermediate Representation) based on the extracted parameters.

[0036] CFG IR includes parameters for each layer required for convolution operations, batch normalization operations, activation and pooling function operations, and parameters for calculating optimized hardware allocation resources (NPU, memory allocation, etc.) required for them.

[0037] Thereafter, the compiler (200) generates a binary code that can control learning and inference operations through the NPU based on the CFG IR and firmware that has data on the hardware-specific optimized resource allocation and distribution process of the NPU embedded device.

[0038] The designed deep learning network model and firmware generated by the compiler (200) are transferred to an embedded device, and deep learning network learning and inference operations are performed.

[0039] There are two main ways to perform hardware-specific optimized resource allocation and distribution processes for embedded devices.

[0040] The first involves grouping input data, dividing it into multiple partial data (Tiling), and then appropriately distributing it to the NPU (140) having an internal systolic array structure. In addition, in order to maximize data reuse, the number of times data such as weight is brought into the On-chip Buffer (141) is reduced as much as possible, thereby minimizing the energy generated in data exchange between the DRAM (120) and the NPU (140). In addition, by identifying the form of the divided input data and designating the number of operating PEs (142), it is necessary to operate using the maximum performance ratio and low power.

[0041] The second is the process of allocating additional memory to store existing memory address access and backpropagation result data so that the NPU, which performs learning as well as inference, can reuse the generated data as much as possible.

[0042] Figure 4 illustrates additional memory allocation for reusing existing data and storing backpropagation data when the NPU performs learning. As illustrated in Figure 4, when the NPU performs a learning task, existing computational data is reused and additional memory addresses for storing backpropagation data are allocated, enabling efficient resource allocation. Since data computed during the forward propagation process is reused during the backpropagation process of learning, the memory address where the data will be loaded is calculated to reuse existing data. In addition, the backpropagation data additionally generated using the data acquired from that address is calculated to be written to the empty space in the already allocated memory, enabling maximum data reuse.

[0043] FIG. 5 is a diagram illustrating the hardware configuration of a deep learning model generation system according to another embodiment of the present invention. The deep learning model generation system according to an embodiment of the present invention can be implemented as a computing system comprising a communication unit (310), an output unit (320), a processor (330), an input unit (340), and a storage unit (350), as illustrated.

[0044] The communication unit (310) is a communication interface for connection with an external network or external device, and is connected to the embedded device. The output unit (320) is an output means for displaying the results of calculations performed by the processor (330), and the input unit (340) is a user interface for receiving user commands and transmitting them to the processor (330).

[0045] The processor (330) generates a pre-designed deep learning network model, executes the compiler illustrated in FIG. 3 described above, generates firmware according to the illustrated procedure, and transmits it to the embedded device via the communication unit (310). The storage unit (350) provides the storage space necessary for the processor (330) to function and operate.

[0046] So far, we have described in detail a preferred embodiment of a method and system for generating a deep learning model for efficient operation of an embedded NPU.

[0047] In an embodiment of the present invention, a method for generating a deep learning model through optimization of data reuse and NPU resource allocation in a compiler that enhances compatibility between deep learning networks designed in various types and embedded NPUs is presented.

[0048] This enables embedded devices with built-in NPUs to secure high compatibility when performing various deep learning network learning and inference operations, and maintain an appropriate balance between NPU performance and energy consumption through optimized computational performance. This also facilitates the development of embedded systems that are tailored to the increasing trend of NPUs that perform both learning and inference.

[0049] Meanwhile, it goes without saying that the technical idea of ​​the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.

[0050] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.

Claims

1. Step to convert deep learning network to ONNX graph; Step of aligning the converted ONNX graph into a directed acyclic graph; Step of extracting parameters for each layer of the sorted ONNX graph; A step of generating CFG IR (Intermediate Representation) based on the extracted parameters; A method for generating a deep learning network model for an embedded device, characterized by comprising the steps of: generating a binary code capable of controlling learning and inference operations through an NPU based on the generated CFG IR, and firmware having data on hardware-specific resource allocation of an NPU embedded device.

2. In claim 1, CFG IR is A method for generating a deep learning network model for an embedded device, characterized in that it includes parameters for calculating layer-by-layer parameters and hardware allocation resources required for convolution operations, batch normalization operations, activation and pooling function operations.

3. In claim 1, A method for generating a deep learning network model for an embedded device, further comprising: a step of transferring a deep learning network and firmware generated in a generation step to an embedded device.

4. In claim 1, Optimized resource allocation for each hardware of NPU embedded devices, A method for creating a deep learning network model for an embedded device, characterized by including a process of grouping input data, dividing the input data into a plurality of partial data, and then distributing the grouped data to an NPU having a systolic array structure.

5. In claim 4, Optimized resource allocation for each hardware of NPU embedded devices, A method for creating a deep learning network model for embedded devices, characterized by minimizing the number of times weights are brought into the internal buffer of an NPU for data reuse.

6. In claim 4, Optimized resource allocation for each hardware of NPU embedded devices, A method for generating a deep learning network model for an embedded device, characterized by specifying the number of PEs from the form of segmented input data.

7. In claim 1, Optimized resource allocation for each hardware of NPU embedded devices, A method for creating a deep learning network model for an embedded device, characterized in that the NPU of the NPU embedded device can access a memory address for data reuse when performing a learning process.

8. In claim 7, The data to be recycled is, A method for creating a deep learning network model for an embedded device, characterized in that the data is calculated in a forward propagation process.

9. In claim 7, Optimized resource allocation for each hardware of NPU embedded devices, A method for creating a deep learning network model for an embedded device, characterized by allocating memory for storing data generated in a backpropagation process performed through data recycling.

10. A processor that converts a deep learning network into an ONNX graph, sorts the converted ONNX graph into a directed acyclic graph, extracts parameters for each layer of the sorted ONNX graph, generates a CFG IR (Intermediate Representation) based on the extracted parameters, and generates a binary code capable of controlling learning and inference operations through the NPU based on the generated CFG IR, and generates firmware with data on hardware-specific resource allocation of the NPU embedded device; and A deep learning network model generation system for an embedded device, characterized in that it includes a storage unit that provides storage space required by a processor.

11. Steps to create a deep learning network and convert the deep learning network into an ONNX graph; A step of generating a firmware having binary code that can control learning and inference operations through NPU from the converted ONNX graph and data on hardware-specific resource allocation of the NPU embedded device; A method for generating a deep learning network model for an embedded device, comprising: a step of transferring a deep learning network and firmware generated in a generation step to an embedded device; 12. A processor that creates a deep learning network, converts the deep learning network into an ONNX graph, and generates binary code that can control learning and inference operations through the NPU from the converted ONNX graph, as well as firmware with data on hardware-specific resource allocation of the NPU embedded device; A deep learning network model generation system for an embedded device, characterized by including a communication unit for transmitting a deep learning network and firmware generated in a generation step to an embedded device.

Citation Information

Patent Citations

  • Runtime optimization of computations of an artificial neural network compiled for execution on a deep learning accelerator

    US20220147813A1