Method for optimizing a neural network model and method for providing a graphical user interface for a neural network model.

The method optimizes neural network models through compression and visualization, addressing inefficiencies in existing methods by enabling efficient optimization and reducing design time through graphical user interfaces.

JP7837774B2Active Publication Date: 2026-03-31SAMSUNG ELECTRONICS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing methods for optimizing neural network models are inefficient and lack effective graphical user interfaces for facilitating model compression and improvement.

Method used

A method involving the compression of pre-trained neural network models and the use of a graphical user interface to visualize and compare original and compressed model information, allowing for efficient optimization and improvement through user interaction.

Benefits of technology

The method enables efficient optimization of neural network models by reducing computational complexity while providing visual guidance for precise adjustments, thereby optimizing network design and reducing design time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837774000001
    Figure 0007837774000001
  • Figure 0007837774000002
    Figure 0007837774000002
  • Figure 0007837774000003
    Figure 0007837774000003
Patent Text Reader

Abstract

To provide a method for optimizing a pre-trained neural network model and a method for providing a graphical user interface for the neural network model.SOLUTION: A method for optimizing a neural network model of the present invention comprises: receiving original model information for a first neural network model pre-trained; generating a second neural network model in which at least part of the first neural network model has been modified by performing compression on the first neural network model and compressed model information for a second neural network model; and visualizing and outputting the compression results such that at least part of the original model information and at least part of the compressed model information are displayed on a single screen.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to machine learning, and more specifically, to a method for optimizing a neural network model and a method for providing a graphical user interface for a neural network model. [Background technology]

[0002] There are various methods for classifying data using machine learning. Among them, data classification methods based on neural networks or artificial neural networks (ANNs) are representative. An artificial neural network is a computational model embodied in software or hardware that mimics the computational power of a biological system using many artificial neurons connected by connecting lines. In an artificial neural network, artificial neurons with simplified functions from biological neurons are used. These are then interconnected through connecting lines with connection strength to replicate human cognitive processes and learning processes.

[0003] In recent years, deep learning techniques have been researched to overcome the limitations of artificial neural networks. As deep learning technology develops, various methods for analyzing and optimizing neural network models have been proposed. For example, some attempts are being made to improve accuracy by providing users with various information about the model, or to provide interfaces to facilitate the reduction of execution time. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2021-34039 [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] The present invention has been made in view of the above-mentioned prior art, and the object of the present invention is to provide a method for efficiently optimizing a pre-trained neural network model.

[0006] Another object of the present invention is to provide a method for providing a graphical user interface for a neural network model. [Means for solving the problem]

[0007] A method for optimizing a neural network model according to one aspect of the present invention, made to achieve the above objective, comprises the steps of: receiving original model information for a pre-trained first neural network model; generating a second neural network model in which at least a part of the first neural network model has been modified by compressing the first neural network model, and compressed model information for the second neural network model; and visualizing and outputting the results of the compression such that at least a part of the original model information and at least a part of the compressed model information are displayed on a single screen.

[0008] A computer-based neural network model processing system according to one embodiment of the present invention comprises an input device, a storage device, an output device, and a processor. The input device receives original model information for a pre-trained first neural network model. The storage device compresses the first neural network model to generate a second neural network model in which at least a portion of the first neural network model has been modified, and compressed model information for the second neural network model. The storage device stores information for program routines that generate the compression results so that at least a portion of the original model information and at least a portion of the compressed model information are displayed on a single screen. The output device visualizes and outputs the compression results. The processor is connected to the input device, the output device, and the storage device to control the execution of the program routines.

[0009] To achieve the above objective, another aspect of the present invention provides a method for optimizing a neural network model, which includes a graphical user interface for optimizing the neural network model.The steps include: providing an interface (GUI); receiving original model information for a first neural network model including a plurality of pre-trained original layers; compressing the first neural network model to modify at least a portion of it and generate a second neural network model including a plurality of compression layers and compressed model information for the second neural network model; displaying a first graphic representation showing the network structure of the plurality of compression layers on the graphic user interface; receiving a first user input from the graphic user interface for the first compression layer among the plurality of compression layers; and displaying a second graphic representation on the graphic user interface to show a comparison of the characteristics of the first original layer corresponding to the first compression layer among the plurality of original layers based on the first user input. The method includes the steps of: displaying on a face; receiving a second user input from the graphic user interface for changing the settings of a second compression layer among the plurality of compression layers; updating the properties of the second compression layer based on the second user input; displaying a third graphic representation on the graphic user interface to show a comparison between the properties of the second original layer corresponding to the second compression layer among the plurality of original layers and the updated properties of the second compression layer based on the second user input; generating a plurality of score values ​​for the plurality of compression layers; displaying a fourth graphic representation on the graphic user interface to show at least a portion of the plurality of compression layers in different ways based on the plurality of score values; and displaying a fifth graphic representation on the graphic user interface to modify at least one of the plurality of compression layers based on the plurality of score values.

[0010] A method for providing a graphic user interface for optimizing a neural network model according to one aspect of the present invention, made to achieve the above objective, comprises the steps of: providing a graphic user interface; receiving first model information for a pre-trained first neural network model; generating a second neural network model in which at least a part of the first neural network model is modified and second model information for the second neural network model by performing data processing on the first neural network model; and displaying a graphic representation on the graphic user interface so as to show a comparison between at least a part of the first model information and at least a part of the second model information.

[0011] An electronic system according to one embodiment of the present invention comprises an input device, a storage device, an output device, and a processor, wherein the input device receives first model information for a pre-trained first neural network model; the storage device generates a second neural network model in which at least a portion of the first neural network model is modified and second model information for the second neural network model by performing data processing on the first neural network model; stores information about a program routine that generates a graphic representation to show a comparison between at least a portion of the first model information and at least a portion of the second model information; the output device provides a graphic user interface to display the graphic representation on the graphic user interface; and the processor is connected to the input device, the storage device, and the output device to control the execution of the program routine. [Effects of the Invention]

[0012] According to the method for optimizing a neural network model of the present invention and the method for providing a graphic user interface therefor, the neural network model can be optimized by performing a compression operation on a pre-trained neural network model instead of a learning operation on the neural network model. As a result, the results before and after compression of the neural network model are compared in terms of layers and / or channels to make them easier to view, the additional information provided becomes information effective for network design and improvement, the information summarized and emphasized by layer grouping enables efficient network development, network design optimized for a specific system becomes possible by providing information to and correcting the target device, and the working time required for model design and improvement can be reduced from the visually provided correction guidelines and the predicted results after correction represented by real-time interaction.

Brief Description of the Drawings

[0013] [Figure 1] It is a sequence diagram showing a method for optimizing a neural network model according to an embodiment of the present invention. [Figure 2] It is a block diagram showing a neural network model processing system according to an embodiment of the present invention. [Figure 3] It is a block diagram showing a neural network model processing system according to an embodiment of the present invention. [Figure 4] It is a block diagram showing a neural network model processing system according to an embodiment of the present invention. [Figure 5a] It is a diagram for explaining a neural network model targeted by a method for optimizing a neural network model according to an embodiment of the present invention. [Figure 5b] It is a diagram for explaining a neural network model targeted by a method for optimizing a neural network model according to an embodiment of the present invention. [Figure 5c]A diagram for explaining a neural network model targeted by an optimization method of a neural network model according to an embodiment of the present invention. [Figure 6] A diagram for explaining a neural network model targeted by an optimization method of a neural network model according to an embodiment of the present invention. [Figure 7] A sequence diagram showing a specific example of an optimization method of the neural network model in FIG. 1. [Figure 8] A sequence diagram showing an example of steps for visualizing and displaying the compression result in FIG. 7. [Figure 9a] A diagram for explaining the operation in FIG. 8. [Figure 9b] A diagram for explaining the operation in FIG. 8. [Figure 10] A sequence diagram showing an example of steps for visualizing and displaying the compression result in FIG. 7. [Figure 11a] A diagram for explaining the operation in FIG. 10. [Figure 11b] A diagram for explaining the operation in FIG. 10. [Figure 11c] A diagram for explaining the operation in FIG. 10. [Figure 12] A sequence diagram showing an example of steps for visualizing and displaying the compression result in FIG. 7. [Figure 13] A diagram for explaining the operation in FIG. 12. [Figure 14] A sequence diagram showing an example of steps for visualizing and displaying the compression result in FIG. 7. [Figure 15a] A diagram for explaining the operation in FIG. 14. [Figure 15b] A diagram for explaining the operation in FIG. 14. [Figure 15c] A diagram for explaining the operation in FIG. 14. [Figure 15d] A diagram for explaining the operation in FIG. 14. [Figure 16] A sequence diagram showing an example of steps for visualizing and displaying the compression result in FIG. 7. [Figure 17] This diagram illustrates the operation shown in Figure 16. [Figure 18] This is a sequence diagram showing an example of the steps for visualizing and displaying the compression results in Figure 7. [Figure 19a] This diagram illustrates the operation shown in Figure 18. [Figure 19b] This diagram illustrates the operation shown in Figure 18. [Figure 19c] This diagram illustrates the operation shown in Figure 18. [Figure 20] This is a sequence diagram illustrating a method for optimizing a neural network model according to one embodiment of the present invention. [Figure 21] Figure 20 is a sequence diagram showing a specific example of an optimization method for the neural network model. [Figure 22] This is a sequence diagram showing an example of the steps to visualize and display the results of the setting changes in Figure 21. [Figure 23a] This diagram illustrates the operation shown in Figure 22. [Figure 23b] This diagram illustrates the operation shown in Figure 22. [Figure 23c] This diagram illustrates the operation shown in Figure 22. [Figure 24] This is a sequence diagram illustrating a method for optimizing a neural network model according to one embodiment of the present invention. [Figure 25] Figure 24 is a sequence diagram showing a specific example of an optimization method for the neural network model. [Figure 26] This is a sequence diagram showing an example of the steps involved in visualizing and displaying the scoring results in Figure 25. [Figure 27a] This diagram illustrates the operation shown in Figure 26. [Figure 27b] This diagram illustrates the operation shown in Figure 26. [Figure 28] This is a sequence diagram showing an example of the steps involved in visualizing and displaying the scoring results in Figure 25. [Figure 29a] This diagram illustrates the operation shown in Figure 28. [Figure 29b] This diagram illustrates the operation shown in Figure 28. [Figure 30] This is a sequence diagram showing a system that embodies the neural network model optimization method according to one embodiment of the present invention. [Modes for carrying out the invention]

[0014] Hereinafter, specific examples of embodiments for carrying out the present invention will be described in detail with reference to the drawings. The same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components will be omitted.

[0015] Figure 1 is a sequence diagram showing a method for optimizing a neural network model according to one embodiment of the present invention.

[0016] As shown in Figure 1, the neural network model optimization method according to this embodiment is performed by a computer-based neural network model processing system, at least in part, implemented in hardware and / or software. The neural network model processing system will be described later with reference to Figures 2 to 4. On the other hand, the neural network model and the neural network system that implements it will be described later with reference to Figures 5 to 6.

[0017] A neural network model optimization method according to one embodiment of the present invention involves receiving original model information (or first model information) for a pre-trained first neural network model (step S100). Compression of the first neural network model is performed to generate a second neural network model in which at least a part of the first neural network model has been modified, and compressed model information (or second model information) for the second neural network model (step S200). However, the present invention is not limited thereto, and the second neural network model and the second model information can also be generated by performing at least one of various data processing, including compression, on the first neural network model.

[0018] The learning process for a neural network model describes the process of how to optimize the solution to a given problem and set of functions, and how to improve the performance and / or accuracy of the neural network model. For example, the learning process for a neural network model includes determining the network structure of the neural network model and determining parameters such as weights. During the learning process for a neural network model, other parameters are changed while maintaining the architecture and data types.

[0019] In contrast, the compression operation on a neural network model refers to the process of reducing the size and computational complexity of a model while maintaining the performance and / or accuracy of a pre-trained neural network model to the greatest extent possible. In other words, increasing the performance and / or accuracy of a neural network model increases the number of layers and parameters, which increases the size and computational complexity of the model. This limits the application of neural network models in environments with limited computation, memory, and power, such as mobile and embedded systems. To overcome these limitations and reduce the complexity of the neural network model, conventional compression operations are performed on pre-trained neural network models. During the compression operation on a neural network model, all parameters, including architecture and data types, are modified.

[0020] As one embodiment, compression techniques for neural network models include quantization, pruning, and matrix decomposition. Quantization is a technique that reduces the storage size of the actual neural network model by reducing the weight values, which are generally represented by floating-point numbers, to a specific number of bits. Pruning is a technique that reduces the size of the model by disconnecting the connections between nodes and weight values ​​that are deemed relatively unnecessary because they have little importance to the model's performance. Matrix decomposition is a technique that reduces the number of weight values ​​and the amount of computation by decomposing one of the weight value matrices of each layer in two or more dimensions into two or more matrices. Examples include low-rank approximation, which decomposes a two-dimensional matrix into two matrices using SVD (singular value decomposition), and canonical polyadic (CP) decomposition, which decomposes a matrix in three or more dimensions into a linear combination of many rank-1 tensors.

[0021] The compression results are visualized and output so that at least a portion of the original model information and at least a portion of the compressed model information are displayed on a single screen (step S300). For example, step S300 is performed using a graphical user interface (GUI). For example, a graphical representation is displayed on the graphical user interface to show a comparison between at least a portion of the first model information and at least a portion of the second model information. The graphical user interface will be described later with reference to Figure 9, etc.

[0022] One embodiment of the present invention provides a method for optimizing a neural network model by performing a compression operation on a pre-trained neural network model, rather than a training operation on the neural network model itself. This method provides a graphical user interface for visually displaying the results of the compression operation and comparing the characteristics before and after the compression operation on a single screen. By providing data on various optimization methods and visually presenting information at a subdivided level, users can perform precise adjustments to their pre-trained neural network model.

[0023] Figures 2, 3, and 4 are block diagrams showing a neural network model processing system according to one embodiment of the present invention.

[0024] As shown in Figure 2, the neural network model processing system 1000 is a computer-based neural network model processing system and includes a processor 1100, a storage device 1200, and an input / output device 1300. The input / output device 1300 includes an input device 1310 and an output device 1320.

[0025] The processor 1100 is used to perform calculations for a neural network model optimization method according to one embodiment of the present invention. For example, the processor 1100 includes a microprocessor, an AP (application processor), a DSP (digital signal processor), a GPU (graphic processing unit), etc. Although only one processor 1100 is shown in Figure 2, the present invention is not limited thereto, and the neural network model processing system 1000 can include multiple processors. On the other hand, although not shown in detail, the processor 1100 includes cache memory to improve computing power.

[0026] The storage device 1200 stores / includes a program (PR) 1210 for a neural network model optimization method according to one embodiment of the present invention, and further stores / includes a compression rule (CR) 1220 and an evaluation rule (ER) 1230 used to perform the neural network model optimization method according to one embodiment of the present invention. The program 1210, the compression rule 1220, and the evaluation rule 1230 are provided from the storage device 1200 to the processor 1100.

[0027] The storage device 1200 includes any storage medium that is computer-readable and stores data and / or instructions issued by a computer. For example, computer-readable storage media include volatile memory such as DRAM (dynamic random access memory), flash memory, MRAM (magnetic random access memory), PRAM (phase change random access memory), and non-volatile memory such as ReRAM (resistance random access memory). The computer-readable storage medium can be inserted into a computer, integrated within a computer, or connected to a computer via a communication medium such as a network and / or wireless link.

[0028] The input device 1310 is used to receive input for a neural network model optimization method according to one embodiment of the present invention. For example, the input device 1310 includes input means for receiving user input (UI), such as a keyboard, keypad, touchpad, touchscreen, mouse, or remote controller.

[0029] The output device 1320 is used to provide an output for a neural network model optimization method according to one embodiment of the present invention. For example, the output device 1320 includes output means for outputting a graphic representation (GR), such as a display device, and further includes other output means such as a speaker or a printer.

[0030] The neural network model processing system 1000 performs a neural network model optimization method according to one embodiment of the present invention described in Figure 1. Specifically, the input device 1310 receives original model information for a pre-trained first neural network model, the storage device 1200 compresses the first neural network model to generate a second neural network model in which at least a part of the first neural network model has been modified, and compressed model information for the second neural network model, and stores information about program routines that generate the compression results so that at least a part of the original model information and at least a part of the compressed model information are displayed on one screen, the output device 1320 visualizes and outputs the compression results, and the processor 1100 is connected to the input device 1310, the storage device 1200, and the output device 1320 to control the execution of the program routines. Furthermore, the neural network model processing system 1000 performs a neural network model optimization method according to one embodiment of the present invention described later in Figures 20 and 24.

[0031] As shown in Figure 3, the neural network model processing system 2000 includes a processor 2100, an input / output device 2200, a network interface 2300, RAM 2400, ROM 2500, and a storage device 2600.

[0032] In one embodiment, the neural network model processing system 2000 is a computer system, and can be a stationary computer system such as a desktop computer, workstation, or server, or a portable computer system such as a laptop computer.

[0033] Processor 2100 is substantially identical to processor 1100 in Figure 2. For example, processor 2100 includes a core capable of executing any instruction set (e.g., IA-32 (Intel Architecture-32), 64-bit extended IA-32, x86-64, PowerPC, Sparc, MIPS, ARM, IA-64, etc.). For example, processor 2100 can access memory, i.e., RAM 2400 or ROM 2500, via a bus and execute instructions stored in RAM 2400 or ROM 2500. As shown in Figure 3, RAM 2400 stores all or part of a program (PR) for a neural network model optimization method according to one embodiment of the present invention, and the program (PR) causes processor 2100 to perform operations for the optimization of the neural network model.

[0034] In other words, a program (PR) includes a set of instructions and / or procedures that can be executed by the processor 2100, and these instructions and / or procedures in the program (PR) cause the processor 2100 to perform actions for optimizing a neural network model according to one embodiment of the present invention. A procedure represents a set of instructions for performing a specific task. Procedures are also called functions, routines, subroutines, subprograms, etc. Each procedure processes data provided from an external source or data generated by another procedure.

[0035] The storage device 2600 is substantially identical to the storage device 1200 in Figure 2. The storage device 2600 stores the program (PR), the compression method (CR), and the evaluation method (ER), and all or part of the program (PR) is loaded from the storage device 2600 into the RAM 2400 before the program (PR) is executed by the processor 2100. The storage device 2600 stores files created in a programming language, and all or part of the program (PR) generated by a compiler or the like is loaded into the RAM 2400.

[0036] The storage device 2600 stores data processed by the processor 2100, or data that has been processed by the processor 2100. That is, the processor 2100 generates new data by processing the data stored in the storage device 2600 according to a program (PR), and stores the generated data in the storage device 2600.

[0037] The input / output device 2200 is substantially identical to the input / output device 1300 in Figure 2. The input / output device 2200 includes input devices such as a keyboard, mouse, and touchscreen, and output devices such as a display device and printer. For example, a user can trigger the execution of a program (PR) by the processor 2100 through the input / output device 2200, input user input (UI) in Figure 2, or view the graphic representation (GR) in Figure 2.

[0038] The network interface 2300 provides the neural network model processing system 2000 with access to an external network. For example, the network includes numerous computer systems and communication links, and the communication links include wired links, optical links, wireless links, or any other form of link. User input (UI) in Figure 2 is provided to the neural network model processing system 2000 via the network interface 2300, and graphic representations (GR) in Figure 2 are provided to other computer systems via the network interface 2300.

[0039] As shown in Figure 4, the neural network model optimization module 100, which is executed and controlled by the neural network model processing systems (1000, 2000) in Figures 2 and 3, includes a compression module 200 and a graphic user interface (GUI) control module 150, and further includes a grouping module 300 and an evaluation and update module 400.

[0040] As used below, "module" refers to software, hardware such as FPGAs or ASICs, or a combination of software and hardware. A "module" is in the form of software and is configured to be stored in an addressable storage medium and executed by one or more processors. For example, a "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. A "module" can also be separated into multiple "modules" that perform specific functions.

[0041] The compression module 200 performs compression operations on the neural network model. For example, the compression module 200 performs compression operations based on the compression scheme (CR in Figures 2 and 3).

[0042] The grouping module 300 performs grouping operations on the layers included in the neural network model. The grouping operations will be described later with reference to Figures 14 and 16.

[0043] The evaluation and update module 400 performs evaluation and update operations (e.g., configuration changes, layer changes, etc.) on the neural network model. For example, the evaluation and update module 400 performs evaluation operations based on the evaluation method (ER in Figures 2 and 3). The evaluation and update operations will be described later with reference to Figures 20 and 24.

[0044] The graphic user interface control module 150 controls the graphic user interface to optimize it for the neural network model. For example, the graphic user interface control module 150 controls the graphic user interface to receive user input (UI in Figure 2) and output a graphic representation (GR in Figure 2).

[0045] In one embodiment, Figures 2 and 3 show the neural network model optimization module 100 as a software entity, but the present invention is not necessarily limited thereto. For example, some or all of the components included in the neural network model optimization module 100 may be embodied in hardware and included in the electronic system of a computer infrastructure.

[0046] Figures 5a to 5c and Figure 6 illustrate the neural network model to which the neural network model optimization method according to one embodiment of the present invention applies.

[0047] Figures 5a to 5c show examples of the network structure of neural network models, and Figure 6 shows an example of a neural network system used to drive a neural network model. For example, a neural network model may include at least one of the following: an Artificial Neural Network (ANN) model, a Convolutional Neural Network (CNN) model, a Recurrent Neural Network (RNN) model, or a Deep Neural Network (DNN) model.

[0048] As shown in Figure 5a, the structure of a typical artificial neural network includes an input layer (IL), multiple hidden layers (HL1, HL2, ..., HLn), and an output layer (OL). The input layer (IL) contains i (where i is a natural number) input nodes (x1, x2, ..., xi), and each input node is input with vector input data (IDAT) of length i.

[0049] Multiple hidden layers (HL1, HL2, ..., HLn) contain n (where n is a natural number) hidden layers and contain hidden nodes (h11, h12, h13, ..., h1m, h21, h22, h23, ..., h2m, hn1, hn2, hn3, ..., hnm). For example, hidden layer (HL1) contains m (where m is a natural number) hidden nodes (h11, h12, h13, ..., h1m), hidden layer (HL2) contains m hidden nodes (h21, h22, h23, ..., h2m), and hidden layer (HLn) contains m hidden nodes (hn1, hn2, hn3, ..., hnm).

[0050] The output layer (OL) contains j (where j is a natural number) output nodes (y1, y2, ..., yj) corresponding to the classes to be classified, and outputs results (e.g., scores or class scores) for each class of the input data (IDAT). The output layer (OL) is also called a fully connected layer and, for example, it numerically represents the probability that the input data (IDAT) corresponds to an automobile.

[0051] The network structure shown in Figure 5a includes connections between nodes (branches) shown by straight lines between two nodes, and weights used in each connection, although these are not shown. Here, nodes within a single layer may not be connected, while nodes in different layers may be fully or partially connected. Each node in Figure 5a (for example, h 1 1) takes the output of the previous node (e.g., x1) as input and performs the calculation, and the result of the calculation is passed to the next node (e.g., h 2 1) Output to this node. Here, each node calculates the output value by applying the input value to a specific function, such as a nonlinear function.

[0052] Generally, the network structure of a neural network is predetermined, and the weights determined by the connections between nodes are calculated using data whose correct classification is already known. This data, whose correct classification is already known, is called "training data," and the process of determining the weights is called "learning." Furthermore, assuming that a set of structures and weights that can be independently learned is called a "model," the process of predicting which class the input data of a model with predetermined weights belongs to and outputting the predicted value is called the "testing" process.

[0053] On the other hand, the general neural network shown in Figure 5a has each node (for example, h 1 1) all nodes of the previous layer (e.g., IL) (e.g., x1, x2, ..., x i ) is linked to the image (IDAT), and when the input data (IDAT) is an image (or audio), the number of required weights increases geometrically as the image size increases, making it unsuitable for handling images. For this reason, convolutional neural networks are being researched, which incorporate filtering techniques into the neural network so that the neural network can learn 2D images well.

[0054] As shown in Figure 5b, the structure of the convolutional neural network includes multiple layers (CONV1, RELU1, CONV2, RELU2, POOL1, CONV3, RELU3, CONV4, RELU4, POOL2, CONV5, RELU5, CONV6, RELU6, POOL3, FC).

[0055] Unlike general neural networks, each layer of a convolutional neural network has three dimensions: width, height, and depth. Consequently, the data input to each layer is also volume data with three dimensions: width, height, and depth. For example, in Figure 5b, if the input image has dimensions of 32 x 32 and three color channels (R, G, B), the input data (IDAT) corresponding to the input image has dimensions of 32 x 32 x 3. The input data (IDAT) in Figure 5b is called input volume data or input activation volume.

[0056] The convolution layers (CONV1, CONV2, CONV3, CONV4, CONV5, CONV6) perform convolution operations on the input. In image processing, convolution refers to processing data using a mask with weighted values, where the input value is multiplied by the mask's weighted values, and the sum of these values ​​is determined as the output value. Here, the mask is also called a filter, window, or kernel.

[0057] Specifically, the parameters of each convolutional layer consist of a series of learnable filters. Each filter is smaller than the total size of the layer in the horizontal and vertical dimensions, but matches the total depth of the layer in the depth dimension. For example, by sliding (more precisely, convolving) each filter along the horizontal and vertical dimensions of the input volume and performing a dot product between the filter and input elements, a two-dimensional activation map is generated. The output volume is generated by stacking such activation maps along the depth dimension. For example, if a convolutional layer (CONV1) applies four filters with zero-padding to an input volume data (IDAT) of size 32×32×3, the output volume of the convolutional layer (CONV1) will have a size of 32×32×12 (i.e., increasing depth).

[0058] The RELU layers (RELU1, RELU2, RELU3, RELU4, RELU5, RELU6) perform corrected linear unit operations on the input. For example, a corrected linear unit operation is a function that treats negative numbers as 0 only, such as max(0, x). For example, if the RELU layer (RELU1) performs a corrected linear unit operation on a 32x32x12 size input volume provided by the convolutional layer (CONV1), the output volume of the RELU layer (RELU1) will have a size of 32x32x12 (i.e., volume maintained).

[0059] The pooling layers (POOL1, POOL2, POOL3) perform downsampling on the horizontal and vertical dimensions of the input volume. For example, when applying a 2x2 filter, it converts the four inputs in a 2x2 region into a single output. Specifically, it either selects the maximum value from the four inputs in a 2x2 region, as in 2x2 maximum pooling, or calculates the average value of the four inputs in a 2x2 region, as in 2x2 average pooling. For example, if the pooling layer (POOL1) applies a 2x2 filter to an input volume of size 32x32x12, the output volume of the pooling layer (POOL1) will have a size of 16x16x12 (i.e., horizontal and vertical reduction, depth maintenance, volume reduction).

[0060] Generally, in a convolutional neural network, one convolutional layer (e.g., CONV1) and one RELU layer (e.g., RELU1) form a pair, and these convolution / RELU layer pairs are repeatedly arranged, with pooling layers inserted in between to reduce the amount of images and extract image features.

[0061] The output layer or fully connected layer (FC) outputs results for each class with respect to the input volume data (IDAT). For example, by repeatedly performing convolution and subsampling, the input volume data (IDAT) corresponding to a 2D image is converted into a 1D matrix (or vector). For example, the fully connected layer (FC) represents the probabilities corresponding to the input volume data (IDAT) being an automobile (CAR), a truck (TRUCK), an airplane (AIRPLANE), a ship (SHIP), a horse (HORSE) as numerical values.

[0062] As shown in FIG. 5c, the structure of the recurrent neural network includes a repetitive structure using the specific nodes (N) or cells shown on the left side of FIG. 5c.

[0063] The structure shown on the right side of FIG. 5c shows the unfolded repetitive connection of the recurrent neural network shown on the left side. To "unfold" a recurrent neural network means to show it for the entire sequence including all nodes (NA, NB, NC). For example, if the sequence information of interest is a sentence consisting of three words, the recurrent neural network is unfolded into a 3-layer neural network structure with one layer per word (without recurrent connections or cycles).

[0064] In a recurrent neural network, X represents the input value of the recurrent neural network. For example, X t is the input value at time step t, and X t-1 and X t+1 are also the input values at time steps t - 1 and t + 1 respectively.

[0065] In a recurrent neural network, S represents the hidden state. For example, S t is the hidden state at time step t, and S t-1 and S t+1These are the hidden states at time steps t-1 and t+1, respectively. The hidden state is calculated from the hidden state value of the previous time step and the input value of the current time step. For example, S t =f(UX t +WS t-1 ) where the nonlinear function f is tanh or ReLU, and S is used to calculate the initial hidden state. -1 It is normally initialized to 0.

[0066] In a regressive neural network, O represents the output value at time step t. For example, O t is the output value at time step t, O t-1 and O t+1 These are the output values ​​at time steps t-1 and t+1, respectively. For example, if we want to predict the next word in a sentence, it will be a probability vector with dimensions equal to the number of words. For example, O t =softscrew(VS t )

[0067] In a regressive neural network, the hidden state is the network's "memory" portion. In other words, a regressive neural network is considered to hold "memory" information for the results calculated up to the present. t It contains all the information about what happened in past time steps, and the output value is O t This regressive neural network relies solely on the memory of the current time step t. Unlike conventional neural network structures where parameter values ​​differ for each hierarchical level, the regressive neural network shares all parameter values ​​(U, V, W in Figure 5c) across all time steps. This indicates that the regressive neural network performs essentially the same calculation for each step, differing only in the input values, thus reducing the number of parameters that need to be learned.

[0068] As shown in Figure 6, the neural network system 500 includes multiple heterogeneous resources for driving the neural network model and a resource management unit 501 for managing and controlling the multiple heterogeneous resources.

[0069] The multiple heterogeneous resources include a CPU 510, an NPU (neural processing unit) 520, a GPU (graphic processing unit) 530, a DSP (digital signal processor) 540, and an ISP (image signal processor) 550, and further include dedicated hardware (DHW) 560, memory (MEM) 570, a DMA (direct memory access) unit 580, and a communication unit 590.

[0070] The CPU 510, NPU 520, GPU 530, DSP 540, ISP 550, and dedicated hardware 560 for specific tasks are referred to as processors, processing units (PEs), computing resources, etc., while the DMA unit 580 and communication unit 590 are referred to as communication resources.

[0071] The CPU 510, NPU 520, GPU 530, DSP 540, ISP 550, and dedicated task hardware 560 perform various functions such as specific calculations or tasks and are used to run neural network models. For example, dedicated task hardware 560 includes a VPU (vision processing unit) and VIP (vision intellectual property). Memory 570 stores data processed by multiple heterogeneous resources and stores data related to the neural network model. The DMA unit 580 controls access to memory 570. For example, the DMA unit 580 includes WDMA (memory DMA), PDMA (peripheral DMA), RDMA (remote DMA), and SDMA (smart DMA). The communication unit 590 performs wired / wireless communication. For example, the communication unit 590 supports internal communications such as the system bus, PCI (peripheral component interconnect), PCIe (PCI express), and / or external communications such as USB (universal serial bus), Ethernet (registered trademark), WiFi, Bluetooth (registered trademark), NFC (near field communication), RFID (radio frequency identification), and mobile telecommunication.

[0072] Although not shown in the diagram, computing resources further include microprocessors, application processors (APs), customized hardware, and compression hardware. Communication resources further include memory copy-capable resources.

[0073] In one embodiment, the neural network system 500 is included in any computer equipment and / or mobile device.

[0074] In one embodiment, various services and / or applications such as computer vision services (e.g., image classification, image detection, image segmentation, image tracking, etc.), biometric authentication services, advanced driver assistance system (ADAS) services, voice assistant services, and automatic speech recognition (ASR) services are executed and processed by the neural network model described in Figures 5a, 5b, and 5c and the neural network system 500 described in Figure 6.

[0075] Figure 7 is a sequence diagram showing a specific example of the optimization method for the neural network model in Figure 1. Explanations that overlap with those in Figure 1 will be omitted below.

[0076] As shown in Figure 7, a neural network model optimization method according to one embodiment of the present invention provides a graphical user interface for optimizing the neural network model (step S500). However, the present invention is not limited thereto, and the graphical user interface is provided to display various information about the neural network model. The specific implementation of the graphical user interface will be described in detail later with reference to Figure 9 and the like.

[0077] The system receives original model information for the pre-trained first neural network model from the graphical user interface (step S100a), compresses the first neural network model to generate the second neural network model and compressed model information for the second neural network model (step S200), and visualizes the compression results and displays them on the graphical user interface so that at least a portion of the original model information and at least a portion of the compressed model information are displayed on one screen (step S300a). For example, as explained in Figures 5a, 5b, and 5c, each of the first and second neural network models is embodied by including multiple layers. The layers included in the first neural network model that correspond to the original model information are defined as multiple original layers, and the layers included in the second neural network model that correspond to the compressed model information are defined as multiple compressed layers.

[0078] Steps S100a and S300a are the same as steps S100 and S300 in Figure 1, and step S200 is substantially the same as step S200 in Figure 1.

[0079] Figure 8 is a sequence diagram showing an example of the steps for visualizing and displaying the compression results from Figure 7. Figures 9a and 9b are diagrams illustrating the operation of Figure 8.

[0080] As shown in Figures 7, 8, 9a, and 9b, the step of visualizing the compression results and displaying them on a graphical user interface (step S300a) involves displaying a graphical representation of the network structure of multiple compression layers included in the second neural network model on the graphical user interface (step S310). For example, the graphical representation shows visual information displayed on one screen of a display device included in the output device 1320 included in the neural network model processing system 1000.

[0081] For example, as shown in Figure 9a, the graphic representation (GR11) shows the network structure of multiple compression layers (LAYER11, LAYER12, LAYER13, LAYER14, LAYER15, LAYER16, LAYER17, LAYER18, LAYER19, LAYER1A, LAYER1B, LAYER1C, LAYER1D, LAYER1E) that exist between the input and output of the second neural network model. The graphic representation (GR11) includes a layer box (e.g., a rectangle) corresponding to each compression layer and arrows indicating the connections between the compression layers.

[0082] In another example, as shown in Figure 9b, the graphic representation (GR12) shows the network structure of multiple compression layers (LAYER11~LAYER1E) and whether each compression layer meets a predetermined criterion. For example, layer boxes corresponding to compression layers that meet the criterion are displayed using the first method, while layer boxes corresponding to compression layers that do not meet the criterion are displayed using a second method different from the first method.

[0083] In one embodiment, as shown in Figure 9b, the first method is a method that shows the layer box without any other indication, and the second method is a method that shows hatching in the layer box. In the example in Figure 9b, the compressed layer (LAYER1E) does not meet the criteria value and is shown with hatching in the layer box, while the remaining compressed layers (LAYER11~LAYER1D) meet the criteria value and are shown in their layer boxes without any other indication. However, the present invention is not limited thereto, and the first and second methods can be embodied using different colors, shapes, etc. For example, the first method is a method that shows a green layer box, and the second method is a method that shows a red layer box.

[0084] In one embodiment, the baseline value relates to a performance (PERF) criterion. For example, the baseline value relates to at least one of various comparison metrics (i.e., indicators for comparing performance), including SQNR (signal-to-quantization-noise power ratio), latency (LTC), power consumption (PWR), and utilization (UTIL). In this case, the layer box is displayed in the first way if the performance value of each compression layer is greater than or equal to the baseline value, and the layer box is displayed in the second way if the performance value of each compression layer is less than the baseline value. In other words, it shows a performance-based index value on a layer-by-layer basis, and if the value is less than a specific value, it can be expressed in another way.

[0085] In one embodiment, the reference value is selectable and / or changeable. For example, as shown in Figure 9b, the reference value is selected and / or changed by selecting at least one of the buttons (112, 114, 116, 118) included in the menu 110 included in the graphic representation (GR12). In the example in Figure 9b, button 112 is selected to select a reference value for SQNR, and each layer box is displayed based on this in one of the first and second methods. For example, the buttons (112, 114, 116, 118) are selected by receiving user input using a mouse, touchscreen, etc., included in the input device 1310 included in the neural network model processing system 1000.

[0086] Figure 10 is a sequence diagram showing an example of the steps for visualizing and displaying the compression results from Figure 7. Figures 11a, 11b, and 11c are diagrams illustrating the operation of Figure 10. Explanations that overlap with those in Figures 8, 9a, and 9b are omitted below.

[0087] As shown in Figures 7, 10, 11a, 11b, and 11c, the step of visualizing the compression results and displaying them on a graphic user interface (step S300a) involves displaying a graphic representation on the graphic user interface that compares the first characteristics for multiple original layers with the second characteristics for multiple compressed layers (step S320). For example, some or all of the first characteristics and some or all of the second characteristics are displayed on a single screen.

[0088] For example, as shown in Figures 11a, 11b, and 11c, each of the graphic representations (GR21, GR22, GR23) compares and displays the distribution characteristics for multiple original layers corresponding to the original model information and the distribution characteristics for multiple compressed layers corresponding to the compressed model information. For example, the original model information is floating model information, and the compressed model information is fixed model information. The distribution characteristics shown on the left correspond to the first characteristics for multiple original layers, and the distribution characteristics shown on the right correspond to the second characteristics for multiple compressed layers. Although not shown in detail, each of the graphic representations (GR21, GR22, GR23) further includes information such as MAC (multiply-accumulate) count values, general operation count values, precision, and performance (e.g., SQNR).

[0089] In one embodiment, the first and second characteristics are represented by at least one of the layer unit and the channel unit (channel-by-channel). In the example in Figures 11a and 11b, the distribution characteristics of multiple channels (channel 10 to channel 23) contained in one original layer and one compressed layer are compared. In the example in Figure 11c, the distribution characteristics of one channel (channel 0) contained in one original layer and one compressed layer are compared.

[0090] In one embodiment, only some data is selectively displayed. For example, in the graphic representations (GR21, GR22) of Figures 11a and 11b, if one channel (channel0) is selected based on user input, the graphic representation (GR23) of Figure 11c is displayed.

[0091] In one embodiment, after a compression operation is performed on the neural network model via a graphical user interface, the output corresponding to the output of the original model is compared and displayed for the modified layer, providing information necessary for model design such as model complexity or capability information, analyzing the model's calculations and attributes to indicate whether they are supported, and providing model memory footprint information from the model analysis.

[0092] Figure 12 is a sequence diagram showing an example of the steps for visualizing and displaying the compression results from Figure 7. Figure 13 is a diagram illustrating the operation of Figure 12. Hereafter, explanations that overlap with Figures 8, 9a, 9b, 10, 11a, 11b, and 11c will be omitted.

[0093] As shown in Figures 7, 12, and 13, in the step of visualizing the compression results and displaying them on a graphical user interface (step S300a), steps S310 and S320 are substantially the same as step S310 in Figure 8 and step S320 in Figure 10, respectively.

[0094] The system receives user input for at least one of the multiple compression layers from the graphical user interface (step S315). Step S320 is performed based on the user input received from step S315.

[0095] For example, step S310 displays one of the graphic representations (GR11, GR12) in Figures 9a and 9b, step S315 receives user input for the first compression layer among multiple compression layers (LAYER11~LAYER1E), and step S320 displays one of the graphic representations (GR21, GR22, GR23) in Figures 11a, 11b, and 11c to show a comparison between the characteristics of the first original layer corresponding to the first compression layer and the characteristics of the first compression layer.

[0096] In one embodiment, as shown in Figure 13, the graphic representation (GRC1) is displayed in a form in which the first graphic representation (GR1) and the second graphic representation (GR2) are combined. For example, the first graphic representation (GR1) corresponds to one of the graphic representations (GR11, GR12) in Figures 9a and 9b, and the second graphic representation (GR2) corresponds to one of the graphic representations (GR21, GR22, GR23) in Figures 11a, 11b, and 11c. In other words, the graphic representations of steps S310 and S320 are displayed on a single screen.

[0097] Figure 14 is a sequence diagram showing an example of the steps for visualizing and displaying the compression results from Figure 7. Figures 15a, 15b, 15c, and 15d are diagrams illustrating the operation of Figure 14. Explanations that overlap with those in Figures 8, 9a, and 9b are omitted below.

[0098] As shown in Figures 7, 14, 15a, 15b, 15c, and 15d, in the step of visualizing the compression results and displaying them on a graphical user interface (step S300a), step S310 is substantially the same as step S310 in Figure 8.

[0099] The system receives user input from the graphical user interface to group multiple compression layers (step S325). A graphical representation of multiple compression layer groups, each containing at least one of the multiple compression layers, is displayed on the graphical user interface (step S330). Step S330 is performed based on the user input received from step S325.

[0100] Layer grouping refers to classifying multiple layers within a neural network model according to specific criteria. When this classification operation is repeated, the entire neural network model can be represented in a reduced form of M layer groups rather than N layers. For example, the number of layer groups is less than or equal to the number of layers (i.e., M <= N). Generally, neural network models consist of tens to hundreds of layers, and the information automatically summarized and highlighted using layer groups compared to the information provided by the layers is efficiently utilized in the development of neural network models.

[0101] In one embodiment, the criterion for grouping multiple compression layers relates to at least one of performance criteria and functional (FUNC) criteria. For example, the performance criteria include at least one of SQNR, latency, power consumption, and usage, while the functional criteria include at least one of CNN, feature extractor, backbone, RNN, LSTM (long short-term memory), and attention module. The criterion changes how the multiple compression layers are grouped.

[0102] For example, as shown in Figure 15a, step S310 displays a graphical representation (GR13) showing the network structure of multiple compression layers (LAYER21, LAYER22, LAYER23, LAYER24, LAYER25, LAYER26) that exist between the input and output of the second neural network model. In other words, the network structure is displayed on a layer-by-layer basis without being grouped.

[0103] As shown in Figures 15a, 15b, 15c, and 15d, the reference value can be selected and / or changed by selecting at least one of the buttons (122, 124, 125, 126, 127, 128) included in the menu 120 included in the graphic representation (GR13) in step S325. As shown in Figures 15b, 15c, and 15d, grouping of multiple compression layers (LAYER21~LAYER26) is automatically performed in step S330, and graphic representations (GR31, GR32, GR33) including the compression layer group are displayed.

[0104] In the example in Figure 15b, buttons (122, 125) are selected to choose a reference value for SQNR, and based on this, a graphic representation (GR31) including compressed layer groups (LAYER_GROUP11, LAYER_GROUP12) and compressed layer (LAYER26) is displayed. For example, compressed layer group (LAYER_GROUP11) includes compressed layers (LAYER21, LAYER22), and compressed layer group (LAYER_GROUP12) includes compressed layers (LAYER23~LAYER25).

[0105] In one embodiment, a compressed layer (LAYER26) in which no compressed layer groups (LAYER_GROUP11, LAYER_GROUP12) are formed in the graphic representation (GR31) indicates that it is a compressed layer that does not meet a predetermined criterion (e.g., a performance-based (SQNR) criterion).

[0106] In the example in Figure 15c, buttons (122, 126) are selected to choose a latency criterion value, and a graphic representation (GR32) including compressed layer groups (LAYER_GROUP21, LAYER_GROUP22, LAYER_GROUP23) is displayed based on this value. For example, compressed layer group (LAYER_GROUP21) includes compressed layers (LAYER21, LAYER22), compressed layer group (LAYER_GROUP22) includes compressed layers (LAYER23, LAYER24), and compressed layer group (LAYER_GROUP23) includes compressed layers (LAYER25, LAYER26). In one embodiment, as described in Figure 15b, compressed layers that do not meet a predetermined criterion (e.g., latency criterion) are displayed layer by layer without being grouped.

[0107] In the example in Figure 15d, buttons (122, 127) are selected to select a baseline value for power consumption, and a graphic representation (GR33) including compressed layer groups (LAYER_GROUP31, LAYER_GROUP32) is displayed based on this value. For example, compressed layer group (LAYER_GROUP31) includes compressed layers (LAYER21~LAYER24), and compressed layer group (LAYER_GROUP32) includes compressed layers (LAYER25, LAYER26). In one embodiment, as described in Figure 15b, compressed layers that do not meet a predetermined standard (e.g., power consumption standard) are displayed individually without being grouped.

[0108] In one embodiment, two or more reference values ​​are applied in a combined form, in which case the result can be displayed in a different form compared to the result of applying only one reference value. For example, when two or more reference values ​​are applied, the highlight layers highlighted by each reference value are displayed in a different way (e.g., in a different color).

[0109] Figure 16 is a sequence diagram showing an example of the steps for visualizing and displaying the compression results from Figure 7. Figure 17 is a diagram illustrating the operation of Figure 16. Explanations that overlap with Figures 8, 9a, 9b, 14, 15a, 15b, 15c, and 15d are omitted below.

[0110] As shown in Figures 7, 16, and 17, in the step of visualizing the compression results and displaying them on a graphical user interface (step S300a), steps (S310, S325, and S330) are substantially the same as steps (S310, S325, and S330) in Figure 14, respectively.

[0111] The system receives user input for at least one of multiple compression layer groups from the graphical user interface (step S335). A graphical representation of the compression layers included in at least one of the compression layer groups is displayed on the graphical user interface (step S340). Step S340 is performed based on the user input received in step S335.

[0112] For example, step S310 displays the graphic representation (GR13) of Figure 15a, steps S325 and S330 display the graphic representation (GR31) of Figure 15b, then step S335 receives user input for the compressed layer group (LAYER_GROUP11), and step S340 displays the graphic representation (GR34) in an expanded form so that the compressed layers (LAYER21, LAYER22) included in the compressed layer group (LAYER_GROUP11) appear.

[0113] In one embodiment, although not shown in detail, after the graphic representation (GR34) of Figure 17 is displayed, user input for the compressed layer group (LAYER_GROUP11) is received again, in which case the graphic representation (GR31) of Figure 15b is displayed again in a reduced form. In other words, the graphic representations (GR31, GR34) of Figures 15b and 17 can be switched between in a form that is expanded or reduced relative to each other.

[0114] Figure 18 is a sequence diagram showing an example of the steps for visualizing and displaying the compression results from Figure 7. Figures 19a, 19b, and 19c are diagrams illustrating the operation of Figure 18. Explanations that overlap with those in Figures 8, 9a, and 9b will be omitted below.

[0115] As shown in Figures 7, 18, 19a, 19b, and 19c, in the step of visualizing the compression results and displaying them on a graphical user interface (step S300a), step S310 is substantially the same as step S310 in Figure 8.

[0116] The system receives user input from the graphical user interface to select at least one target device that will execute multiple compression layers (step S345). For example, the target device includes at least one of the CPU 510, NPU 520, GPU 530, DSP 540, and ISP 550 in Figure 6, and at least one additional resource. A graphical representation indicating whether the multiple compression layers are suitable for the target device is displayed on the graphical user interface (step S350). Step S350 is performed based on the user input received in step S345.

[0117] For example, as shown in Figure 19a, step S310 displays a graphic representation (GR14) showing the network structure of multiple compression layers (LAYER31, LAYER32, LAYER33, LAYER34, LAYER35, LAYER36) that exist between the input and output of the second neural network model.

[0118] As shown in Figures 19a, 19b, and 19c, step S345 allows the target device to be selected and / or modified by selecting at least one of the buttons (132, 134, 136) included in menu 130, which is included in the graphic representation (GR14). As shown in Figures 19b and 19c, step S350 displays graphic representations (GR41, GR42) indicating whether the multiple compression layers (LAYER31~LAYER36) are suitable for the selected target device.

[0119] In the example shown in Figure 19b, button 132 is selected, and the NPU is selected as the target device. This displays a graphic representation (GR41) indicating whether multiple compression layers (LAYER31~LAYER36) are suitable for being driven by the NPU.

[0120] In the example in Figure 19c, buttons (132, 136) are selected, and the NPU and DSP are selected as the target devices. This displays a graphic representation (GR42) indicating whether multiple compression layers (LAYER31~LAYER36) are suitable for being driven by the NPU and DSP.

[0121] In one embodiment, a compression layer deleted or removed from the graphic representation (GR41, GR42) (for example, the compression layer (LAYER32) in Figure 19b) indicates that it is a compression layer that cannot be driven by the target device (i.e., the NPU). Also, a compression layer displayed in a different manner from other compression layers in the graphic representation (GR41, GR42) (for example, the hatched compression layers (LAYER33, LAYER35) in Figures 19b and 19c) indicates that it is a compression layer unsuitable for the target device.

[0122] In one embodiment, modifications to layers that are undriveable or unsuitable by the target device can be proposed. For example, the target device can automatically modify layers to optimize the model's performance and display the modified layers or the modifiable layers. For instance, the modification of layers can be proposed using a modification method determined by the target device, and the processing time of the selected layer can be predicted and proposed by the target device through reinforcement learning or the like. In this case, by modifying the neural network model by the selected target device, the neural network model is modified to suit the device and / or system that intends to use it, and the modified model can be easily compared by the user. As another example, it is also possible to propose modifying the target device.

[0123] However, the present invention is not limited thereto, and not only layer changes but also layer group changes can be proposed and implemented.

[0124] On the other hand, in the embodiments, in the examples of Figures 14, 16, and 18, step S320 of Figure 10 and / or steps (S315, S320) of Figure 12 are further performed, and in the examples of Figures 14 and 16, steps (S325, S330) of Figure 14 and / or steps (S325, S330, S335, S340) of Figure 16 are further performed.

[0125] Figure 20 is a sequence diagram showing a method for optimizing a neural network model according to one embodiment of the present invention. Explanations that overlap with those in Figure 1 are omitted below.

[0126] As shown in Figure 20, in the neural network model optimization method according to one embodiment of the present invention, the steps (S100, S200, and S300) are substantially the same as the steps (S100, S200, and S300) in Figure 1.

[0127] The settings of the second neural network model are changed to improve its performance, and the results of the setting changes are visualized and output (step S600). For example, similar to step S300, step S600 is performed using a graphical user interface.

[0128] Figure 21 is a sequence diagram showing a specific example of the optimization method for the neural network model in Figure 20. Explanations that overlap with those in Figures 7 and 20 will be omitted below.

[0129] As shown in Figure 21, in the neural network model optimization method according to one embodiment of the present invention, steps (S500, S100a, S200, and S300a) are substantially the same as steps (S500, S100a, S200, and S300a) in Figure 7.

[0130] The settings of the second neural network model are changed to improve its performance, and the results of the setting changes are visualized and displayed on the graphical user interface (step S600a). Step S600a is the same as step S600 in Figure 20.

[0131] Figure 22 is a sequence diagram showing an example of the steps to visualize and display the results of the setting changes in Figure 21. Figures 23a, 23b, and 23c are diagrams to explain the operation of Figure 22. Hereafter, explanations that overlap with Figures 8, 9a, 9b, 10, 11a, 11b, 11c, 12, and 13 will be omitted.

[0132] As shown in Figures 21, 22, 23a, 23b, and 23c, the step of visualizing the results of the setting changes and displaying them on the graphical user interface (step S600a) is to receive user input for setting changes for multiple compression layers from the graphical user interface (step S605). The second characteristics for the multiple compression layers are updated (step S610). Step S610 is performed based on the user input received from step S605.

[0133] A graphical representation comparing the first characteristics for multiple original layers with the updated second characteristics for multiple compressed layers is displayed on the graphical user interface (step S620). Step S620 is the same as step S320 in Figure 10.

[0134] For example, as shown in Figure 23a, a graphic representation (GRC21) combining the first graphic representation (GR15) and the second graphic representation (GR24) is displayed before step S600a is performed. The first graphic representation (GR15) shows the network structure of multiple compression layers (LAYER41, LAYER42, LAYER43, LAYER44, LAYER45, LAYER46) that exist between the input and output of the second neural network model, and the second graphic representation (GR24) shows a comparison of the distribution characteristics of multiple original layers corresponding to the original model information and the distribution characteristics of multiple compression layers corresponding to the compressed model information.

[0135] As shown in Figure 23b, step S605 modifies the settings of the compression layer (LAYER42) based on user input from menu 140 included in the graphic representation (GRC22) which is a combination of the first graphic representation (GR51) and the second graphic representation (GR24), thereby forming / providing a modified compression layer (LAYER42'). For example, the number of bits (BN) of the input and / or output of the compression layer (LAYER42) can be changed from X to Y (where X and Y are integers of 1 or more).

[0136] In one embodiment, the setting change is performed when it is determined that the performance of the second neural network model obtained as a result of the compression operation is lower than the performance of the first neural network model before the compression operation. For example, as shown in the second graphic representation (GR24), the setting change is performed when the distribution characteristics for the compressed layer are worse than the distribution characteristics for the original layer.

[0137] As shown in Figure 23c, step S610 performs a characteristic update operation, step S620 displays a second graphic representation (GR52) comparing the distribution characteristics for multiple original layers with the updated distribution characteristics for multiple compressed layers, and a graphic representation (GRC23) combining the first graphic representation (GR51) and the second graphic representation (GR52) is displayed. For example, when using a modified compressed layer (LAYER42'), the updated distribution characteristics for the compressed layer are better than the distribution characteristics for the original layer.

[0138] As mentioned above, real-time interaction allows for immediate application and verification of improvements to model performance and their effects. In other words, users can be shown the information they need through real-time interaction. For example, necessary information includes feature-map distribution, SQNR, SNR, MAC count, OP count, etc. This shortens development time, allows for more detailed results to be verified, and enables users to freely check the expected performance of device-specific models while designing them.

[0139] Figure 24 is a sequence diagram showing a method for optimizing a neural network model according to one embodiment of the present invention. Explanations that overlap with those in Figure 1 are omitted below.

[0140] As shown in Figure 24, in the neural network model optimization method according to one embodiment of the present invention, steps (S100, S200, and S300) are substantially the same as steps (S100, S200, and S300) in Figure 1.

[0141] Step S700 performs scoring to determine the operational efficiency of the second neural network model, and visualizes and outputs the scoring results. For example, similar to step S300, step S700 is performed using a graphical user interface.

[0142] Figure 25 is a sequence diagram showing a specific example of the optimization method for the neural network model in Figure 24. Explanations that overlap with those in Figures 7 and 24 will be omitted below.

[0143] As shown in Figure 25, in the neural network model optimization method according to one embodiment of the present invention, steps (S500, S100a, S200, and S300a) are substantially the same as steps (S500, S100a, S200, and S300a) in Figure 7.

[0144] The second neural network model is scored to determine its operational efficiency, and the scoring results are visualized and displayed on the graphical user interface (step S700a). Step S700a is the same as step S700 in Figure 24.

[0145] Figure 26 is a sequence diagram showing an example of the steps for visualizing and displaying the scoring results from Figure 25. Figures 27a and 27b are diagrams illustrating the operation of Figure 26. Explanations that overlap with Figures 8, 9a, and 9b are omitted below.

[0146] As shown in Figures 25, 26, 27a, and 27b, the step of visualizing the scoring results and displaying them on a graphical user interface (step S700a) involves generating multiple score values ​​for multiple compression layers (step S710), and displaying graphical representations on the graphical user interface that show at least some of the multiple compression layers in different ways based on the multiple score values ​​(step S720).

[0147] For example, as shown in Figure 27a, a first graphic representation (GR16) is displayed showing the network structure of multiple compression layers (LAYER51, LAYER52, LAYER53, LAYER54, LAYER55, LAYER56) that exist between the input and output of the second neural network model before step S700a is performed.

[0148] As shown in Figure 27b, step S710 is performed to generate multiple score values ​​(SV51, SV52, SV53, SV54, SV55, SV56) for multiple compression layers (LAYER51 to LAYER56), and step S720 is performed to display graphic representations (GR61) that show some of the compression layers (LAYER54 to LAYER56) in different ways based on the multiple score values ​​(SV51 to SV56). In one embodiment, the multiple score values ​​(SV51 to SV56) are also displayed.

[0149] In one embodiment, layer boxes corresponding to compression layers with a score value greater than a reference score value are displayed in a first manner, while layer boxes corresponding to compression layers with a score value less than or equal to the reference score value are displayed in a second manner different from the first manner.

[0150] In one embodiment, as shown in Figure 27b, the first method is a method that shows the layer box without any other display, and the second method is a method that displays hatching on the layer box. In the example in Figure 27b, the hatched compressed layers (LAYER54~LAYER56) indicate layers with relatively lower operational efficiency, where the narrower the hatching interval, the lower the operational efficiency of the layer. However, the present invention is not limited thereto, and for example, the darker the color in which the layer box is displayed, the lower the operational efficiency of the layer.

[0151] In one embodiment, the score value includes at least one of the following: the performance estimation result for each layer, whether or not the structure / layer type is unsuitable for the device characteristics and the resulting score, the capacity prediction value for each layer, and the memory footprint usage. For example, the above indicators may be calculated by summing them up using different weighting values.

[0152] A neural network model is formed by combining layers with various characteristics and structures where several types of layers are grouped together. Each layer or model structure may or may not be efficient for the operation of a particular device and / or system. As mentioned above, scoring can detect inefficient layers and model structures, and by displaying them and providing a user-modifiable interface, improved performance for optimized modeling can be achieved.

[0153] Figure 28 is a sequence diagram showing an example of the steps for visualizing and displaying the scoring results from Figure 25. Figures 29a and 29b are diagrams illustrating the operation of Figure 28. Explanations that overlap with those in Figures 8, 9a, 9b, 26, 27a, and 27b will be omitted below.

[0154] As shown in Figures 25, 28, 29a, and 29b, in the step of visualizing the scoring results and displaying them on a graphic user interface (step S700a), steps (S710 and S720) are substantially the same as steps (S710 and S720) in Figure 26, respectively.

[0155] Based on the scoring results, modify at least one of the compression layers (step S730).

[0156] For example, as shown in Figure 29a, when step S730 is performed and one of the buttons (152, 154) included in menu 150 included in graphic representation (GR62) is selected, compression layer (LAYER61) is selected from compression layers (LAYER61, LAYER62), and the compression layer with the lowest operational efficiency (LAYER56) is changed to compression layer (LAYER61). In the example in Figure 22, only the compression layer settings are changed, but in the example in Figure 28, the compression layer itself can be changed.

[0157] As shown in Figure 29b, step S730 is performed and a graphic representation (GR63) including the compressed layers (LAYER51~LAYER55, LAYER61) and score values ​​(SV51~SV55, SV61) is displayed. The hatching interval of the compressed layer (LAYER61) is wider than that of the compressed layer (LAYER56), confirming that the operational efficiency has improved.

[0158] In one embodiment, a layer or region containing multiple layers is selected by dragging. In another embodiment, a more suitable layer or structure is recommended for the selected layer or region, and one can be selected from the recommended list, or a layer to be modified can be selected from a layer palette containing various layers. A graphical representation of the modification results is displayed as shown in Figure 29b.

[0159] According to embodiments of the present invention, it is possible to provide a GUI that visually displays the compression results of a neural network model and modifies parameters on a layer-by-layer basis, a tool that compares and visualizes the results of the neural network model before and after compression, a tool that visualizes indicators that serve as criteria for evaluating the compression results, a tool that matches information modified after compression with the original information, a tool that reconstructs and visualizes the network graph as needed, a tool that shows modifiable layers for each target device and proposes modification methods, and an interface that shows and modifies suggested improvements and the necessary information for model design and improvement, as well as a tool that shows the expected improvement performance in real time.

[0160] Figure 30 is a sequence diagram showing a system in which a neural network model optimization method according to one embodiment of the present invention is implemented.

[0161] As shown in Figure 30, the system 3000 includes a user device 3100, a cloud computing environment 3200, and a network 3300. The user device 3100 includes a neural network model optimization engine 3110, and the cloud computing environment 3200 includes cloud storage 3210, a database 3220, a neural network model optimization engine 3230, a cloud neural network model engine 3240, and an inventory 3250. The neural network model optimization method according to this embodiment is implemented on the cloud environment, and the neural network model optimization method is performed by the neural network model optimization engines (3110, 3230). [Industrial applicability]

[0162] Embodiments of the present invention are applicable to various devices and systems that embody artificial neural networks and / or machine learning. For example, embodiments of the present invention are more usefully applicable to electronic systems such as PCs, server computers, data centers, workstations, laptops, mobile phones, smartphones, MP3 players, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), digital TVs, digital cameras, portable game consoles, navigation devices, wearable devices, IoT (Internet of Things) devices, IoE (Internet of Everything) devices, e-books, VR (Virtual Reality) devices, AR (Augmented Reality) devices, and drones.

[0163] Although embodiments of the present invention have been described in detail above with reference to the drawings, the present invention is not limited to the embodiments described above, and can be modified and implemented in various ways without departing from the technical scope of the present invention. [Explanation of Symbols]

[0164] 100 Neural Network Model Optimization Modules 110, 120, 130, 140 menu Buttons 112, 114, 116, 118, 122, 124, 125, 126, 127, 128, 132, 134, 136, 152, 154 150 Graphical User Interface (GUI) Control Modules 200 compression modules 300 Grouping Modules 400 evaluation and update modules 500 Neural Network Systems 501 Resource Management Department 510 CPU 520 NPU(neural processing unit) 530 GPU(graphic processing unit)530 540 DSP(digital signal processor) 550 ISP(image signal processor) 560 Dedicated hardware (DHW) for specific tasks 570 memory 580 DMA (direct memory access) section 590 Communications Department 1000, 2000 Neural Network Model Processing Systems 1100, 2100 processors 1200, 2600 storage device 1210 Program (PR) 1220 Compression rule (CR) 1230 Evaluation Rule (ER) 1300, 2200 Input / Output Devices 1310 Input device 1320 Output device 2300 Network Interfaces 2400 RAM 2500 ROM 3000 System 3100 User device 3110, 3230 Neural Network Model Optimization Engine 3200 Cloud Computing Environments 3210 Cloud Storage 3220 Databases 3240 Cloud Neural Network Model Engine 3250 Inventory 3300 Network

Claims

1. The steps include receiving original model information for the first neural network model that was pre-trained, The steps include generating a second neural network model in which at least a part of the first neural network model is modified by compressing the first neural network model, and generating compressed model information for the second neural network model, The process includes the step of visualizing and outputting the compression result such that at least a portion of the original model information and at least a portion of the compressed model information are displayed on a single screen. The first neural network model described above includes multiple original layers, The second neural network model includes multiple compression layers, A method for optimizing a neural network model, characterized in that the step of visualizing and outputting the compression results includes displaying a first graphic representation on a graphic user interface that shows a comparison between a first characteristic of the plurality of original layers and a second characteristic of the plurality of compressed layers.

2. The method further includes the step of providing a graphical user interface for optimizing the first neural network model, The neural network model optimization method according to claim 1, characterized in that at least a portion of the original model information and at least a portion of the compressed model information are displayed on the graphic user interface.

3. The method for optimizing a neural network model according to claim 1, characterized in that the first characteristic and the second characteristic are represented by at least one of a layer unit and a channel unit.

4. The step of visualizing and outputting the results of the compression is as follows: The steps include displaying a second graphic representation showing the network structure of the plurality of compression layers on the graphic user interface, The process further includes the step of receiving a first user input for a first compression layer among the plurality of compression layers from the graphic user interface, The method for optimizing a neural network model according to claim 1, characterized in that the first graphic representation is displayed to show a comparison between the characteristics of the first original layer corresponding to the first compressed layer among the plurality of original layers based on the first user input.

5. The second graphic representation includes a plurality of layer boxes corresponding to the plurality of compression layers, The first layer box corresponding to a compression layer among the plurality of compression layers that satisfies a predetermined standard value is displayed in the first manner. The neural network model optimization method according to claim 4, characterized in that the second layer box corresponding to a compression layer among the plurality of compression layers that does not meet the criteria value is displayed in a second manner different from the first manner.

6. The method for optimizing a neural network model according to claim 5, characterized in that the reference value is related to at least one of a plurality of comparison metrics, including SQNR, latency, power consumption, and usage.

7. The step of visualizing and outputting the results of the compression is as follows: The steps include receiving a second user input from the graphic user interface for grouping the plurality of compression layers, The method for optimizing a neural network model according to claim 4, further comprising the step of displaying a third graphic representation on the graphic user interface, each representing a plurality of compression layer groups, each containing at least one of the plurality of compression layers, based on the second user input.

8. The step of visualizing and outputting the results of the compression is as follows: The steps include receiving a third user input from the graphic user interface for a first compression layer group among the plurality of compression layer groups, The method for optimizing a neural network model according to claim 7, further comprising the step of displaying a fourth graphic representation on the graphic user interface that shows the compression layers included in the first compression layer group based on the third user input.

9. The step of visualizing and outputting the results of the compression is as follows: The steps include receiving a second user input from the graphic user interface for selecting at least one target device that performs the plurality of compression layers, The method for optimizing a neural network model according to claim 4, further comprising the step of displaying a third graphic representation on the graphic user interface indicating whether the plurality of compression layers are suitable for the target device based on the second user input.

10. The neural network model optimization method according to claim 9, characterized in that the target device includes at least one of a CPU, NPU, GPU, DSP, and ISP.

11. The neural network model optimization method according to claim 1, further comprising the step of making setting changes to improve the performance of the second neural network model, and visualizing and outputting the results of the setting changes.

12. The step of visualizing and outputting the results of the aforementioned setting changes is: The steps include receiving a first user input from the graphic user interface for changing the settings for the plurality of compression layers, A step of updating the second characteristic based on the first user input, A method for optimizing a neural network model according to claim 11, comprising the step of displaying a second graphic representation on the graphic user interface that shows a comparison between the first characteristic and the updated second characteristic.

13. The method for optimizing a neural network model according to claim 1, further comprising the step of performing scoring to determine the operational efficiency of the second neural network model, and visualizing and outputting the results of the scoring.

14. The step of visualizing and outputting the results of the aforementioned scoring is: The steps include generating multiple score values ​​for the multiple compression layers, A method for optimizing a neural network model according to claim 13, comprising the step of displaying a second graphic representation on the graphic user interface that shows at least a portion of the plurality of compression layers in a different manner based on the plurality of score values.

15. The second graphic representation includes a plurality of layer boxes corresponding to the plurality of compression layers, The first layer box corresponding to the compression layer among the plurality of compression layers whose score value is greater than the reference score value is displayed in the first manner. The neural network model optimization method according to claim 14, characterized in that a second layer box corresponding to a compression layer whose score value among the plurality of compression layers is smaller than the reference score value or the same compression layer is displayed in a second manner different from the first manner.

16. The method for optimizing a neural network model according to claim 14, characterized in that the plurality of score values ​​are obtained based on at least one of the following: the compression performance estimation result for the plurality of compression layers, whether or not it is suitable for the device characteristics, the layer type, the capacity prediction result, and the memory footprint usage.

17. The method for optimizing a neural network model according to claim 13, further comprising the step of modifying at least one of the plurality of compression layers based on the results of the scoring.

18. The steps include providing a graphical user interface for optimizing a neural network model, A step of receiving original model information for a first neural network model that includes multiple pre-trained original layers, The steps include: compressing the first neural network model thereby modifying at least a portion of the first neural network model to generate a second neural network model including multiple compression layers and compressed model information for the second neural network model; The steps include displaying a first graphic representation showing the network structure of the plurality of compression layers on the graphic user interface, The steps include receiving a first user input for a first compression layer among the plurality of compression layers from the graphic user interface, The steps include displaying a second graphic representation on the graphic user interface so as to show a comparison between the characteristics of the first original layer corresponding to the first compressed layer among the plurality of original layers based on the first user input, and the characteristics of the first compressed layer; The steps include receiving a second user input from the graphic user interface for changing the settings of a second compression layer among the plurality of compression layers, The steps include updating the characteristics of the second compression layer based on the second user input, The steps of displaying a third graphic representation on the graphic user interface so as to show a comparison between the characteristics of the second original layer corresponding to the second compressed layer among the plurality of original layers based on the second user input, and the characteristics of the updated second compressed layer, The steps include generating multiple score values ​​for the multiple compression layers, The steps include displaying a fourth graphic representation on the graphic user interface that shows at least a portion of the plurality of compression layers in a different manner based on the plurality of score values, A method for optimizing a neural network model, comprising the step of displaying a fifth graphic representation on the graphic user interface so as to modify at least one of the plurality of compression layers based on the plurality of score values.

19. A method for providing a graphical user interface for optimizing a neural network model, Steps include providing a graphical user interface, A step of receiving first model information for a pre-trained first neural network model, The steps include generating a second neural network model in which at least a part of the first neural network model is modified by performing data processing on the first neural network model, and generating second model information for the second neural network model, The method includes the step of displaying a graphic representation on the graphic user interface so as to show a comparison between at least a portion of the first model information and at least a portion of the second model information, The first neural network model described above includes multiple original layers, The second neural network model includes multiple compression layers, The method is characterized in that the step of displaying the graphic representation on the graphic user interface includes the step of displaying a first graphic representation on the graphic user interface that shows a comparison between a first characteristic of the plurality of original layers and a second characteristic of the plurality of compressed layers.

Citation Information

Patent Citations

  • Structure display method for neural network

    JP1992190461A

  • Neural network method and apparatus

    JP2021034039A

  • Learning device, learning system, and learning method

    JP2021039640A

  • No-coding machine learning pipeline

    US20210055915A1

  • Information processing method and information processing device

    WO2017141517A1