Apparatus and method for quantizing neural network model

CN120035831APending Publication Date: 2025-05-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072788.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-26
Filing Date
2023-11-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Deep learning is difficult to achieve high performance in embedded systems with limited hardware resources, mainly due to the need for large operations and large memory capacity.

Method used

An electronic device and a control method are provided. By obtaining neural network model, test data and performance condition information, quantifying at least one layer of a plurality of layers in the neural network model, generating a first quantized neural network model, and sending it to a target device for testing, determining whether the performance of the quantized model meets the required conditions based on the result data and configuration file information, and adding it to the quantized model candidate list if it is satisfied.

Benefits of technology

By quantifying neural network models, the demand for hardware resources is reduced and the performance and efficiency of deep learning models in embedded systems with limited hardware resources is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035831A_ABST
    Figure CN120035831A_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device includes a memory and at least one processor, and the at least one processor: obtains information about a neural network model and test data and desired performance conditions for the neural network model; obtaining a first quantized neural network model by quantizing at least one of a plurality of layers included in the neural network model; sending the first quantitative neural network model and the test data to a target device; receiving, from the target device, result data obtained from the first quantized neural network model using the test data as an input, and profile information of the target device associated with the first quantized neural network model; and based on the result data and the configuration file information, if the first quantitative neural network model meets the required performance condition, adding the first quantitative neural network model to an available quantitative neural network model candidate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an electronic device and a control method thereof. More specifically, the present disclosure relates to an electronic device of a quantized neural network model and a control method thereof. Background Art

[0002] Recently, artificial intelligence systems that achieve human-level intelligence are being developed. An artificial intelligence system refers to a system in which a machine makes learning and determinations on its own, unlike the rule-based intelligent system of the prior art, and it is used in various fields such as speech recognition, image recognition, and future prediction. More specifically, recently, artificial intelligence systems that solve given problems based on deep learning are being developed.

[0003] The application of deep learning has shown very high performance in speech and image recognition, natural language processing, etc. Therefore, the demand for on-device deep learning for real-life artificial intelligence services is also increasing.

[0004] However, deep learning requires a large number of operations and large memory capacity, making it difficult to achieve high performance in embedded systems with limited hardware resources.

[0005] To solve this problem, lightweight deep learning models with low operation complexity are being developed. As one of the lightweight deep learning model technologies, there is a quantization technology that limits the values ​​that can be represented by the weights of the deep learning model.

[0006] The above information is presented as background information only to assist in understanding the present disclosure. No determination has been made, and no statement is made, as to whether any of the above may be applicable as prior art with respect to the present disclosure. Summary of the invention

[0007] Aspects of the present disclosure are to at least address the above-mentioned problems and / or disadvantages and to provide at least the advantages described below. Therefore, one aspect of the present disclosure is to provide an apparatus and method for quantizing a neural network model.

[0008] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments.

[0009] According to one aspect of the present disclosure, an electronic device is provided. The electronic device includes a memory and at least one processor, wherein the at least one processor is configured to: obtain a neural network model, test data for the neural network model, and information about required performance conditions; quantize at least one layer of a plurality of layers included in the neural network model, and obtain a first quantized neural network model; send the first quantized neural network model and the test data to a target device; receive from the target device result data obtained from the first quantized neural network model with the test data as input, and profile information of the target device related to the first quantized neural network model; and based on the result data and the profile information, add the first quantized neural network model to available quantized neural network model candidates based on the first quantized neural network model satisfying the required performance conditions.

[0010] According to another aspect of the present disclosure, a method for controlling an electronic device is provided. The method includes: obtaining a neural network model, test data for the neural network model, and information about required performance conditions; quantizing at least one layer of a plurality of layers included in the neural network model, and obtaining a first quantized neural network model; sending the first quantized neural network model and the test data to a target device; receiving from the target device result data obtained from the first quantized neural network model with the test data as input, and profile information of the target device related to the first quantized neural network model; and based on the result data and the profile information, adding the first quantized neural network model to available quantized neural network model candidates based on the first quantized neural network model satisfying the required performance conditions.

[0011] According to another aspect of the present disclosure, one or more non-transitory computer-readable storage media storing computer-executable instructions are provided. When the computer-executable instructions are executed by at least one processor of an electronic device, the electronic device is configured to perform operations, wherein the operations include: obtaining a neural network model, test data for the neural network model, and information about required performance conditions; quantizing at least one layer of a plurality of layers included in the neural network model, and obtaining a first quantized neural network model; sending the first quantized neural network model and the test data to a target device; receiving from the target device result data obtained from the first quantized neural network model with the test data as input, and profile information of the target device related to the first quantized neural network model; and based on the result data and the profile information, adding the first quantized neural network model to available quantized neural network model candidates based on the first quantized neural network model satisfying the required performance conditions.

[0012] Other aspects, advantages, and salient features of the disclosure will become apparent to those skilled in the art from the following detailed description, which, taken in conjunction with the annexed drawings, discloses various embodiments of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings, in which:

[0014] Figure 1 is a block diagram for illustrating a configuration of an electronic device according to an embodiment of the present disclosure;

[0015] Figure 2 is a diagram for illustrating the operation of an electronic device according to an embodiment of the present disclosure;

[0016] Figure 3 is a flowchart for illustrating the operation of an electronic device according to an embodiment of the present disclosure;

[0017] Figure 4 is a diagram for illustrating a method for an electronic device to acquire a quantized neural network model according to an embodiment of the present disclosure;

[0018] Figure 5 is a diagram for illustrating a method for an electronic device to obtain profile information and a quantization error of a target device according to an embodiment of the present disclosure;

[0019] Figure 6 is a flowchart for illustrating a method for an electronic device to obtain profile information and a quantization error of a target device according to an embodiment of the present disclosure;

[0020] Figure 7 is a timing diagram for illustrating operations of an electronic device and a target device according to an embodiment of the present disclosure; and

[0021] Figure 8 is a flowchart for illustrating a control method of an electronic device according to an embodiment of the present disclosure.

[0022] The same reference numerals are used throughout the drawings to denote the same elements. DETAILED DESCRIPTION

[0023] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details should be considered as merely exemplary. Therefore, it will be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0024] The terms and words used in the following description and claims are not limited to the bibliographical meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0025] It should be understood that singular forms include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces. Furthermore, in describing the present disclosure, in cases where it is determined that a detailed explanation of a related known function or feature may unnecessarily obscure the subject matter of the present disclosure, the detailed explanation will be omitted.

[0026] In addition, the embodiments described below can be modified in various forms, and the scope of the technical ideas of the present disclosure is not limited to the following embodiments. On the contrary, these embodiments are provided to make the present disclosure more complete and complete, and to fully convey the technical ideas of the present disclosure to those skilled in the art.

[0027] In addition, the terms used in the present disclosure are only used to describe specific embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In addition, a singular expression includes a plural expression unless it is clearly defined differently in the context.

[0028] Furthermore, in the present disclosure, expressions such as “having”, “may have”, “including”, and “may include” indicate the presence of such characteristics (for example, elements such as numbers, functions, operations, and components), and do not exclude the presence of additional characteristics.

[0029] In addition, in the present disclosure, the expression "A or B", "at least one of A and / or B", or "one or more of A and / or B", etc. may include all possible combinations of the listed items. For example, "A or B", "at least one of A and B", or "at least one of A or B" may refer to all of the following situations: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B.

[0030] In addition, the expressions "first", "second", etc. used in the present disclosure may describe various elements regardless of any order and / or importance. In addition, such expressions are only used to distinguish one element from another element and are not intended to limit the elements.

[0031] In addition, the description in the present disclosure that one element (e.g., a first element) is “(operably or communicatively) coupled” to another element (e.g., a second element) and / or is “(operably or communicatively) coupled to” another element (e.g., the second element), or is “connected to” another element (e.g., the second element) should be interpreted as including the case where the one element is directly coupled to the other element and the case where the one element is coupled to the other element via another element (e.g., a third element).

[0032] Conversely, description that one element (e.g., a first element) is “directly coupled” or “directly connected” to another element (e.g., a second element) may be interpreted as indicating that there is no further element (e.g., a third element) between the one element and the other element.

[0033] Furthermore, the expression "configured to" used in the present disclosure may be used interchangeably with other expressions such as "suitable for," "capable of," "designed to," "suitable for," "manufactured to," and "capable of," depending on the circumstances. Furthermore, the term "configured to" does not necessarily mean that a device is "specifically designed to" in terms of hardware.

[0034] Conversely, in some cases, the expression "a device configured to..." may mean that the device is "capable" of performing an operation together with another device or component. For example, the phrase "a processor configured to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing the corresponding operations, or a general-purpose processor (e.g., a central processing unit (CPU) or an application processor) that can perform the corresponding operations by executing one or more software programs stored in a memory device.

[0035] In addition, in the embodiments of the present disclosure, a "module" or "component" may perform at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. In addition, in addition to a "module" or "component" that needs to be implemented as specific hardware, multiple "modules" or "components" may be integrated into at least one module and implemented as at least one processor.

[0036] In addition, various elements and regions in the drawings are schematically shown. Therefore, the technical concept of the present disclosure is not limited by the relative sizes or intervals shown in the drawings.

[0037] Hereinafter, embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art to which the present disclosure pertains can easily implement the present disclosure.

[0038] Figure 1 is a block diagram for illustrating a configuration of an electronic device according to an embodiment of the present disclosure.

[0039] refer to Figure 1 , the electronic device 100 may include a memory 110, a communication interface 120, and a processor 130. In the electronic device 100, some of the above components may be omitted, or other components may be further included.

[0040] In addition, the electronic device 100 may be implemented as a server, but this is merely an example, and the electronic device 100 may be implemented in various forms such as a smart phone, a television (TV), a smart TV, a set-top box, a mobile phone, a personal digital assistant (PDA), a laptop computer, a media player, an e-book terminal, a terminal for digital broadcasting, a navigation, a self-service terminal, an MP3 player, a wearable device, a home appliance, and other mobile or non-mobile computing devices.

[0041] The memory 110 may store at least one instruction regarding the electronic device 100. The memory 110 may also store an operating system (O / S) for operating the electronic device 100. In addition, the memory 110 may store various software programs or applications for the electronic device 100 to operate according to various embodiments of the present disclosure. In addition, the memory 110 may include a semiconductor memory (such as a flash memory) or a magnetic storage medium (such as a hard disk), etc.

[0042] Specifically, the memory 110 may store various software modules for the electronic device 100 to operate according to various embodiments of the present disclosure, and the processor 130 may control the operation of the electronic device 100 by executing the various software modules stored in the memory 110. For example, the memory 110 may be accessed by the processor 130, and the processor 130 may perform reading / recording / correction / deletion / update of data, etc.

[0043] In addition, in the present disclosure, the term memory 110 may be used as a meaning including the memory 110, a read-only memory (ROM) (not shown) and a random access memory (RAM) (not shown) inside the processor 130, or a memory card (not shown) (e.g., a micro secure digital (SD) card, a memory stick) installed on the electronic device 100.

[0044] In addition, the communication interface 120 includes a circuit and is a component that can communicate with an external device and a server. The communication interface 120 can perform communication with an external device or a server based on a wired or wireless communication method. The communication interface 120 may include a Bluetooth module (not shown), a Wi-Fi module (not shown), an infrared (IR) module, a local area network (LAN) module, an Ethernet module, etc. Here, each communication module can be implemented in the form of at least one hardware chip. The wireless communication module may include at least one communication chip that performs communication according to various wireless communication protocols (such as Zigbee, Universal Serial Bus (USB), Mobile Industry Processor Interface Camera Serial Interface (MIPI CSI), third generation (3G), third generation partnership project (3GPP), long term evolution (LTE), LTE advanced (LTE-A), fourth generation (4G), fifth generation (5G), etc.) in addition to the communication method mentioned above. However, these are only examples, and the communication interface 120 can use at least one communication module in various communication modules.

[0045] The processor 130 may control the overall operation and functions of the electronic device 100. Specifically, the processor 130 may be connected to components of the electronic device 100 including the memory 110, and as described above, may control the overall operation of the electronic device 100 by executing at least one instruction stored in the memory 110.

[0046] The processor 130 may be implemented in various ways. For example, the processor 130 may be implemented as at least one of an application specific integrated circuit (ASIC), a logic integrated circuit, an embedded processor, a microprocessor, a hardware control logic, a hardware finite state machine (FSM), or a digital signal processor (DSP). In addition, in the present disclosure, the term processor 130 may be used as a meaning including a central processing unit (CPU), a graphics processing unit (GPU), a main processing unit (MPU), etc.

[0047] More specifically, the processor 130 may include at least one processor. Specifically, the at least one processor may include one or more of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an integrated many-core (MIC), a digital signal processor (DSP), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. At least one processor may control one or a random combination of other components of the electronic device and perform operations related to communication or data processing. In addition, at least one processor may execute one or more programs or instructions stored in a memory. For example, at least one processor may execute a method according to one or more embodiments of the present disclosure by executing at least one instruction stored in a memory.

[0048] In the case where the method according to one or more embodiments of the present disclosure includes multiple operations, the multiple operations may be performed by one processor, or by multiple processors. For example, when the first operation, the second operation, and the third operation are performed by the method according to one or more embodiments of the present disclosure, all of the first operation, the second operation, and the third operation may be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor), and the third operation may be performed by the second processor (e.g., an artificial intelligence dedicated processor).

[0049] At least one processor may be implemented as a single-core processor including one core, or it may be implemented as one or more multi-core processors including multiple cores (e.g., multiple cores of the same type or multiple cores of different types). In the case where at least one processor is implemented as a multi-core processor, each of the multiple cores included in the multi-core processor may include an internal memory of the processor (such as a cache memory, an on-chip memory, etc.), and a public cache shared by the multiple cores may be included in the multi-core processor. In addition, each of the multiple cores included in the multi-core processor (or some of the multiple cores) may independently read program instructions for implementing the method according to one or more embodiments of the present disclosure and execute the instructions, or all of the multiple cores (or some of the cores) may be linked to each other, and read program instructions for implementing the method according to one or more embodiments of the present disclosure and execute the instructions.

[0050] In the case where the method according to one or more embodiments of the present disclosure includes multiple operations, the multiple operations may be performed by one of the multiple cores included in the multi-core processor, or they may be implemented by multiple cores. For example, when the first operation, the second operation, and the third operation are performed by the method according to one or more embodiments of the present disclosure, all of the first operation, the second operation, and the third operation may be performed by the first core included in the multi-core processor, or the first operation and the second operation may be performed by the first core included in the multi-core processor, and the third operation may be performed by the second core included in the multi-core processor.

[0051] In an embodiment of the present disclosure, the processor 130 may represent a system on chip (SoC) in which at least one processor and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor. In addition, here, the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, a machine learning accelerator, etc., but the embodiments of the present disclosure are not limited thereto.

[0052] The operations of the processor 130 for implementing various embodiments of the present disclosure may be implemented through a plurality of modules.

[0053] Specifically, data about the plurality of modules according to the present disclosure may be stored in the memory 110, and the processor 130 may access the memory 110 and load the data about the plurality of modules into a memory or a buffer within the processor 130, and then implement various embodiments according to the present disclosure by using the plurality of modules. Here, the plurality of modules may include a neural network model acquisition module 131, a quantization module 132, a profile acquisition module 133, a quantization error identification module 134, and a neural network model selection module 135.

[0054] Furthermore, at least one of the plurality of modules according to the present disclosure may be implemented as hardware and included within the processor 130 in the form of a system on chip.

[0055] Alternatively, at least one of the plurality of modules according to the present disclosure may be implemented as a separate external device, and the electronic device 100 and each module may perform the operation according to the present disclosure while performing communication.

[0056] Hereinafter, the operation of the electronic device 100 according to the present disclosure will be described in detail with reference to the accompanying drawings.

[0057] Figure 2 is a diagram for illustrating the operation of an electronic device according to an embodiment of the present disclosure. Figure 2 The neural network model acquisition module 131 may acquire at least one of the neural network model 10, test data 20 about the neural network model, or information 30 about required performance conditions.

[0058] Here, the neural network model acquisition module 131 may acquire at least one of the neural network model 10 , test data 20 about the neural network model, or information 30 about required performance conditions from an external server, an external device, or a target device 200 through the communication interface 120 .

[0059] Optionally, at least one of the neural network model 10 , the test data 20 regarding the neural network model 20 , or the information 30 regarding the required performance conditions may be stored in the memory 110 .

[0060] Here, the neural network model 10 may be a pre-trained neural network model. In addition, here, each weight of the multiple layers included in the neural network model 10 may be a floating point number. For example, each weight may be represented by a floating point 32 format.

[0061] In addition, the test data 20 about the neural network model may represent test data for evaluating the performance of the neural network model. Specifically, the test data 20 may include data for obtaining result data of the neural network model 10 or the quantized neural network model. Here, the result data may represent data output by the neural network model 10 or the quantized neural network model when the test data 20 is input into the neural network model 10 or the quantized neural network model.

[0062] In addition, the required performance condition may represent at least one condition that the performance of the quantized neural network model in the target device 200 should satisfy. Here, the required performance condition may represent a performance condition required for a specific application. For example, the required performance condition may include at least one of the inference delay time caused when the inference of the neural network model is performed in the target device 200, the memory usage and power consumption of the target device 200, the accuracy of the quantized neural network model, the quantization error between the quantized neural network model and the neural network model 10, or the size of the quantized neural network model.

[0063] The quantization module 132 may quantize at least one layer among the plurality of layers included in the neural network model 10 using the first bit precision and obtain a quantized neural network model.

[0064] When the quantized neural network model is acquired, the configuration file acquisition module 133 may send the quantized neural network model and the test data 20 to the target device 200 .

[0065] Thereafter, the target device 200 may perform inference of the quantized neural network model by using the received quantized neural network model and the test data 20 .

[0066] When receiving the quantized neural network model and the test data 20, the target device 200 may perform inference of the received quantized neural network model and obtain result data of the quantized neural network model and profile information of the target device 200 related to the quantized neural network model.

[0067] Here, the result data of the quantized neural network model may represent data output by the quantized neural network model when the test data 20 is respectively input into each quantized neural network model. For example, if the quantized neural network model is a model for classifying a category, the result data may include a probability value of the predicted category.

[0068] The profile information related to the quantized neural network model of the target device 200 may include information about at least one of the time incurred for inference (i.e., inference latency), the memory usage incurred for inference, or the power consumption incurred for inference when the target device 200 performs inference of the quantized neural network model.

[0069] Specifically, the target device 200 may perform inference of the quantized neural network model by using the test data 20. For example, the target device 200 may input the test data 20 into the quantized neural network model.

[0070] Here, the target device 200 can perform inference of the quantized neural network model by using the test data 20 and obtain result data output by the quantized neural network model. The target device 200 can perform inference of the quantized neural network model by using the test data 20 and obtain profile information of the target device 200 related to the quantized neural network model.

[0071] When the result data and the profile information are obtained, the target device 200 may send the result data and the profile information to the electronic device 100. For example, the profile acquisition module 133 may acquire the result data obtained from the quantized neural network model and the profile information of the target device 200 related to the quantized neural network model.

[0072] Then, by using the acquired result data, the quantization error identification module 134 may identify a quantization error between the neural network model 10 and the quantized neural network model.

[0073] Specifically, the quantization error identification module 134 may compare result data of the quantized neural network model obtained by using the test data 20 with result data of the neural network model 10 obtained by using the test data 20, and identify a quantization error between the neural network model 10 and the quantized neural network model.

[0074] When a quantization error is identified, the neural network model selection module 135 may identify whether the quantized neural network model meets the required performance conditions. If the quantized neural network model meets the required performance conditions, the neural network model selection module 135 may add the quantized neural network model to the candidate 40 of available quantized neural network models.

[0075] In addition, if the quantized neural network model does not meet the required performance conditions, the neural network model selection module 135 may quantize the neural network model 10 using a second bit precision based on information about a quantization method for quantizing the quantized neural network model, and obtain a neural network model quantized using the second bit precision.

[0076] Subsequently, the plurality of modules may repeatedly perform the aforementioned operations and quantize the neural network model 10, and update the available quantized neural network model candidates 40 by using the quantized neural network model.

[0077] The operation of the electronic device 100 including a plurality of modules for quantizing the neural network model 10 and updating the available candidates 40 of the quantized neural network model by using the quantized neural network model will be described through the following drawings.

[0078] Figure 3 is a flowchart for illustrating the operation of the electronic device according to an embodiment of the present disclosure.

[0079] refer to Figure 3 In operation S305, the processor 130 may quantize at least one of the multiple layers included in the neural network model 10 using the nth bit precision. Here, the processor 130 may quantize the neural network model 10 using the nth bit precision and obtain a first quantized neural network model. Specifically, the processor 130 may quantize the neural network model 10 using the nth bit precision and obtain a plurality of first quantized neural network models including the first quantized neural network model. Here, n may be a natural number.

[0080] Here, if n is 1, the processor 130 may quantize the neural network model 10 using the first bit precision and obtain the neural network model quantized using the first bit precision. Here, the processor 130 may quantize at least one layer among the multiple layers included in the neural network model 10 using the first bit precision and quantize the remaining layers among the multiple layers except the at least one layer using the lowest bit precision.

[0081] The first bit precision may indicate a bit precision one level higher than the lowest bit precision. Here, information about the lowest bit precision may be stored in the memory 110.

[0082] In addition, when n is 2 or greater, a method for the processor 130 to quantize the neural network model 10 may be the same as the method to be described below in operation S350.

[0083] Specifically, the processor 130 may quantize a weight included in at least one layer among a plurality of layers included in the neural network model 10 and an intermediate value output by the at least one layer using the nth bit.

[0084] The processor 130 may quantize weights included in remaining layers except for at least one layer among the plurality of layers included in the neural network model 10 and intermediate values ​​output by the remaining layers using the lowest bit precision.

[0085] Here, the processor 130 may change a layer to be quantized among a plurality of layers included in the neural network model 10, and obtain a plurality of first quantized neural network models including a first quantized neural network model.

[0086] For example, the processor 130 may select a layer in the neural network model 10 to be quantized using a first bit precision, and quantize the selected layer using the first bit precision. For example, the lowest bit precision may be four-bit precision, and the first bit precision may be five-bit precision. Here, the processor 130 may quantize at least one layer of the multiple layers using five-bit precision, and quantize the remaining layers except for at least one layer of the multiple layers using four-bit precision. For example, in the case where three layers are included in the neural network model 10, the processor 130 may quantize each layer in the order of [5 bits, 4 bits, 4 bits], [4 bits, 5 bits, 4 bits] or [4 bits, 4 bits, 5 bits], [5 bits, 5 bits, 4 bits], [5 bits, 4 bits, 5 bits], [4 bits, 5 bits, 5 bits] or [5 bits, 5 bits, 5 bits], and obtain multiple first quantized neural network models.

[0087] refer to Figure 4 The processor 130 may select the first layer in the neural network model 10 as the layer to be quantized, and quantize the first layer in the neural network model 10 with five-bit precision, and quantize the remaining layers with four-bit precision, and obtain the first quantized neural network model 11.

[0088] Optionally, the processor 130 may select the second layer in the neural network model 10 as the layer to be quantized, and quantize the second layer in the neural network model 10 with five-bit precision, and quantize the remaining layers with four-bit precision, and obtain the first quantized neural network model 12.

[0089] For example, the processor 130 may change a layer in the neural network model 10 that is to be quantized using the first bit precision and obtain a plurality of first quantized layers.

[0090] Thereafter, the processor 130 may transmit the quantized neural network model and the test data 20 to the target device 200 at operation S310 .

[0091] Specifically, when the neural network model 10 is quantized, the processor 130 may send the first quantized neural network model and the test data 20 to the target device 200. Here, the processor 130 may send a plurality of first quantized neural network models including the first quantized neural network model and the test data 20 to the target device 200. Here, in the event that there is a history that the test data 20 was previously sent to the target device 200, the processor 130 may exclude the test data 20 and send a plurality of first quantized neural network models including the first quantized neural network model to the target device 200.

[0092] Subsequently, the target device 200 may perform inference of the first quantitative neural network model by using the received first quantitative neural network model and the test data 20. In addition, the target device 200 may perform inference of the first quantitative neural network model, and obtain result data of the first quantitative neural network model and profile information of the target device 200 for the first quantitative neural network model. Subsequently, the target device 200 may send the obtained result data and the profile information of the target device 200 to the electronic device 100. For example, the electronic device 100 may obtain, from the target device 200, the result data obtained from the first neural network model with the test data 20 as input, and profile information of the target device 200 related to the first quantitative neural network model.

[0093] Specifically, the target device 200 may perform inference of each of the plurality of first quantized neural network models including the first quantized neural network model by using the test data 20, and obtain result data of each of the plurality of first quantized neural network models and profile information of the target device 200 related to each of the plurality of first quantized neural network models. Subsequently, the target device 200 may send the result data obtained from each of the plurality of first neural network models with the test data as input, and the profile information of the target device 200 related to each of the plurality of first neural network models to the electronic device 100.

[0094] For example, in operation S315, the electronic device 100 may obtain, from the target device 200, result data obtained from each of the plurality of first neural network models with the test data as input, and profile information of the target device 200 related to the plurality of first quantized neural network models.

[0095] Here, the result data of each of the plurality of first quantitative neural network models including the first quantitative neural network model may represent data output by each of the plurality of first quantitative neural network models when the test data 20 is input to each of the plurality of first quantitative neural network models. For example, if each of the plurality of first quantitative neural network models is a model for classifying a category, the result data may represent a probability value of the predicted category.

[0096] The profile information of the target device 200 associated with each of the multiple first quantized neural network models including the first quantized neural network model may include information about at least one of the time incurred for inference (i.e., inference latency), the memory usage incurred for inference, or the power consumption incurred for inference when the target device 200 performs inference of each of the multiple first quantized neural network models.

[0097] Specifically, the target device 200 may perform inference of each of the plurality of first quantized neural network models by using the test data 20. For example, the target device 200 may input the test data 20 into each of the plurality of first quantized neural network models.

[0098] Here, the target device 200 may perform inference of each of the plurality of first quantized neural network models including the first quantized neural network model by using the test data 20, and obtain result data output by each of the plurality of first quantized neural network models. In addition, the target device 200 may perform inference of each of the plurality of first quantized neural network models by using the test data 20, and obtain profile information of the target device 200 related to each of the plurality of first quantized neural network models.

[0099] In addition, in the present disclosure, the electronic device 100 may send a plurality of first quantized neural network models including a first quantized neural network model and test data to the target device 200, and receive result data and profile information of each of the plurality of first quantized neural network models including the first quantized neural network model from the target device 200. However, this is merely an example, and the electronic device 100 may perform inference of each of the plurality of first quantized neural network models including the first quantized neural network model by using a software simulator included in the electronic device 100. Here, the software simulator may be a program installed in the memory 110, but this is merely an example, and the software simulator may be a separate module included in the electronic device 100. The electronic device 100 may obtain result data and profile information of each of at least one neural network model from the software simulator included in the electronic device 100. Here, the result data and profile information obtained by using the software simulator may be the same as the result data and profile information received from the target device 200.

[0100] Subsequently, in operation S320, the processor 130 may identify a quantization error between the neural network model 10 and the quantized neural network model by using the acquired result data.

[0101] Specifically, the processor 130 may identify a quantization error between the neural network model 10 and each of a plurality of first quantization neural network models including the first quantization neural network model by using result data acquired from the target device 200 .

[0102] Specifically, the processor 130 may compare each of the result data of each of the multiple first quantized neural network models including the first quantized neural network model obtained by using the test data 20 with the result data of the neural network model 10, and identify the quantization error between the neural network model 10 and each of the multiple first quantized neural network models.

[0103] Subsequently, in operation S325 , the processor 130 may add a neural network model that satisfies a required performance condition among the quantized neural network models to the candidates 40 of available quantized neural network models.

[0104] Specifically, the processor 130 may identify whether each of the plurality of first quantized neural network models including the first quantized neural network model satisfies the required performance condition. Subsequently, the processor 130 may add the neural network model that satisfies the required performance condition among the plurality of first quantized neural network models to the candidate 40 of the available quantized neural network models.

[0105] Here, if the first quantized neural network model satisfies the required performance conditions, the processor 130 may add the first quantized neural network model to the available quantized neural network model candidates 40. In addition, if the first quantized neural network model does not satisfy the required performance conditions, and the number of neural network models included in the available quantized neural network model candidates 40 is less than a predetermined number, the processor 130 may quantize the neural network model 10 using a second bit precision and obtain a second quantized neural network model.

[0106] Specifically, the processor 130 may identify whether each of the plurality of first quantized neural network models satisfies the required performance condition. Here, the processor 130 may identify whether the profile information of the target device 200 related to each of the plurality of first quantized neural network models satisfies the required performance condition.

[0107] Subsequently, the processor 130 may identify a neural network model in which the profile information of the target device 200 satisfies the required performance condition among the plurality of first quantized neural network models.

[0108] Subsequently, the processor 130 may add a neural network model that satisfies the required performance condition among the plurality of first quantized neural network models to the candidates 40 of available quantized neural network models.

[0109] Subsequently, in operation S330 , the processor 130 may identify whether the number of neural network models included in the candidates 40 of available quantization neural network models is less than a predetermined number k1.

[0110] Here, if the number of neural network models included in the candidate 40 of the available quantized neural network model is less than the predetermined number k1 (in operation S330-Yes), then in operation S340, based on the profile information and the quantization error, the processor 130 may obtain a score for each of the remaining neural network models (i.e., the neural network models that do not satisfy the required performance condition among the plurality of first quantized neural network models) other than the neural network model included in the candidate 40 of the available quantized neural network model. In the present disclosure, the score may be a score indicating suitability for the required performance.

[0111] Specifically, the processor 130 may obtain a score for each condition included in the required performance conditions for the quantized neural network model.

[0112] For example, the processor 130 may obtain a score regarding inference latency (latency-aware score) for a specific quantized neural network model by using the following Equation 1.

[0113] Equation 1

[0114]

[0115] Here, qerror_lowest may represent a quantization error between a model of the neural network model 10 that is quantized using the lowest bit precision and the neural network model 10. In addition, qerror_quantization may represent a quantization error between a specific quantized neural network model and the neural network model 10. In addition, latency_quantization may represent an inference delay when the target device 200 performs inference by using a specific quantized neural network model. In addition, latency_lowest may represent an inference delay when the target device 200 performs inference by using a model of the neural network model 10 that is quantized using the lowest bit precision.

[0116] The processor 130 may obtain a score regarding the model size (a model size-aware score) for a specific quantized neural network model by using the following Equation 2.

[0117] Equation 2

[0118]

[0119] Here, size_quantization may represent the size of a specific quantized neural network model. In addition, size_lowest may represent the size of a model that is a neural network model 10 quantized using the lowest bit precision.

[0120] In addition, the processor 130 may obtain a score (runtime memory awareness score) regarding memory usage of the target device 200 when the target device 200 performs inference by using a specific quantized neural network model by using the following Equation 3.

[0121] Equation 3

[0122]

[0123] Here, memory_quantization may indicate the memory usage of the target device 200 when the target device 200 performs inference by using a specific quantized neural network model. In addition, memory_lowest may indicate the memory usage of the target device 200 when the target device 200 performs inference by using a model that is a neural network model 10 quantized using the lowest bit precision.

[0124] In addition, the processor 130 may obtain a score (power consumption aware score) regarding power consumption of the target device 200 when the target device 200 performs inference by using a specific quantized neural network model by using Equation 4 below.

[0125] Equation 4

[0126]

[0127] Here, power_quantization may represent the power consumption of the target device 200 when the target device 200 performs inference by using a specific quantized neural network model. In addition, power_lowest may represent the power consumption of the target device 200 when the target device 200 performs inference by using a model that is the neural network model 10 quantized using the lowest bit precision.

[0128] Then, the processor 130 may identify a weighted sum of the scores regarding each condition included in the required performance conditions as a score for the specific quantized neural network model.

[0129] For example, processor 130 may identify a score for a particular quantized neural network model by using Equation 5 below.

[0130] Equation 5

[0131] Score=aS1+bS2+cS3+dS4

[0132] Here, a, b, c, and d may be random real numbers. In addition, S1, S2, S3, and S4 may be scores regarding each condition included in the required performance condition. For example, S1 may represent a score regarding inference latency, S2 may represent a score regarding model size, S3 may represent a score regarding memory usage, and S4 may represent a score regarding power consumption.

[0133] Subsequently, in operation S345, the processor 130 may identify a predetermined number k2 of neural network models in order of having higher scores among the remaining neural network models except for the neural network model included in the candidates 40 of the available quantized neural network models among the plurality of first quantized neural network models.

[0134] Specifically, based on the acquired scores, the processor 130 may identify a predetermined number k2 of neural network models in order of higher scores among the neural network models in the plurality of first quantized neural network models except for the neural network models included in the candidates 40 of the available quantized neural network models.

[0135] In addition, here, the processor 130 may identify a predetermined number k2 of neural network models based on the acquired scores and certain constraints.

[0136] Specifically, the processor 130 may identify a neural network model that satisfies a specific restriction condition among the neural network models in the plurality of first quantized neural network models except the neural network models included in the candidates 40 of the available quantized neural network models. Subsequently, the processor 130 may identify a predetermined number k2 of neural network models in order of higher scores among the neural network models that satisfy the specific restriction condition.

[0137] Here, the specific restriction condition may be a restriction condition of the profile information of the target device 200. For example, the specific restriction condition may include at least one of the following conditions: a condition that the inference delay time included in the profile information of the target device 200 is less than or equal to a predetermined value, a condition that the memory usage caused for inference is less than or equal to a predetermined value, or a condition that the power consumption caused for inference is less than or equal to a predetermined value.

[0138] For example, if the specific restriction condition is a condition that the inference delay is less than or equal to a predetermined value, the processor 130 may identify a predetermined number k2 of neural network models in order of having higher scores among the neural network models having an inference delay less than or equal to the predetermined value among the neural network models among the multiple first quantized neural network models excluding the neural network models included in the candidates 40 of the available quantized neural network models.

[0139] Furthermore, in the present disclosure, the predetermined numbers k1 and k2 may be the same, but this is merely an example, and k1 and k2 may be different.

[0140] In addition, the predetermined numbers k1 and k2 may be fixed integers, but this is merely an example, and the numbers may vary according to the performance of the electronic device 100 or the number of layers included in the neural network model. For example, the predetermined numbers k1 and k2 may be determined by at least one of the layers included in the neural network model 10 or the performance of the electronic device 100. Here, information about the predetermined numbers k1 and k2 may be pre-stored in the memory 110, but this is merely an example, and the processor 130 may identify the predetermined number k1 or k2 based on the performance of the electronic device 100 or the number of layers included in the neural network model 10.

[0141] For example, if the performance of the electronic device 100 is the first performance, the processor 130 may identify m1 as the predetermined number k1 or k2. In addition, if the performance of the electronic device 100 is the second performance higher than the first performance, the processor 130 may identify m2 higher than m1 as the predetermined number k1 or k2.

[0142] Alternatively, if the number of layers included in the neural network model 10 is p1, the processor 130 may identify m1 as the predetermined number k1 or k2. In addition, if the number of layers included in the neural network model 10 is p2 higher than p1, the processor 130 may identify m2 higher than m1 as the predetermined number k1 or k2.

[0143] In addition, according to an embodiment of the present disclosure, the processor 130 may identify a predetermined number k2 of quantized neural network models in order of higher scores. However, this is merely an example, and the processor 130 may identify a predetermined number k2 of quantized neural network models by comparing a threshold value with the score of the quantized neural network model.

[0144] Specifically, the processor 130 may identify the number k2 of quantized neural network models by using the statistics of the obtained scores. Here, the processor 130 may identify the standard deviation of the scores obtained from each of the neural network models other than the models included in the candidates 40 of the available quantized neural network models among the plurality of first quantized neural network models. Alternatively, the processor 130 may identify the standard deviation of the scores obtained from each of the neural network models satisfying the restriction conditions among the neural network models not included in the candidates 40 of the available quantized neural network models among the plurality of first quantized neural network models.

[0145] Then, when the standard deviation is identified, the processor 130 may identify a threshold value by using Equation 6 below.

[0146] Equation 6

[0147] Threshold = n × standard deviation

[0148] Here, n can be a random constant.

[0149] When the threshold is identified, the processor 130 may identify a neural network model whose score is greater than or equal to the identified threshold among the neural network models other than the models included in the candidates 40 of the available quantized neural network models among the plurality of first quantized neural network models. Alternatively, the processor 130 may identify a neural network model whose score is greater than or equal to the identified threshold among the neural network models satisfying the restriction conditions among the neural network models other than the models included in the candidates 40 of the available quantized neural network models among the plurality of first quantized neural network models. Here, the number of identified neural network models may be k2.

[0150] For example, the number of identified neural network models may be dynamically changed based on statistics that quantify the scores of the neural network models.

[0151] Thereafter, in operation S350, based on information about a predetermined number k2 of neural network models, the processor 130 may quantize at least one layer of the plurality of layers included in the neural network model 10 using an (n+1)th bit precision and obtain a second quantized neural network model.

[0152] Specifically, based on information about a quantization method for quantizing a predetermined number k2 of neural network models, the processor 130 can quantize at least one layer of the multiple layers included in the neural network model 10 using the (n+1)th bit precision and obtain a second quantized neural network model.

[0153] Here, the processor 130 may quantize at least one of the multiple layers included in the neural network model 10 using the (n+1)th bit precision, and obtain multiple second quantized neural network models including a second quantized neural network model. Here, the (n+1)th bit precision may represent a bit precision one level higher than the nth bit precision.

[0154] Specifically, the processor 130 may quantize at least one layer quantized using the lowest bit precision in one model of the predetermined number k2 of neural network models using the (n+1)th bit precision. Subsequently, the processor 130 may quantize the layers other than the at least one layer among the multiple layers using the same bit precision as the bit precision of the one model.

[0155] Here, if n is 1, based on information about a method for quantizing a predetermined number k2 of neural network models, the processor 130 may quantize at least one of the multiple layers included in the neural network model 10 using a second bit precision, and obtain multiple second quantized neural network models including a second quantized neural network model.

[0156] Specifically, the processor 130 may quantize at least one of the layers quantized using the lowest bit precision in one of the predetermined number k2 of neural network models among the multiple layers using the second bit precision. Subsequently, among the multiple layers included in one of the predetermined number k2 of neural network models, the processor 130 may quantize the remaining layers except the at least one layer using the same bit precision as the bit precision of the one of the predetermined number k2 of neural network models.

[0157] For example, if the first quantized neural network model does not meet the required performance condition, the processor 130 may quantize at least one of the multiple layers included in the neural network model 10 using the second bit precision, and obtain multiple second quantized neural network models including the second quantized neural network model. Figure 4 , the first quantized neural network model may be model 11, in which three layers are quantized in the order of [5 bits, 4 bits, 4 bits]. Here, the layers quantized using the lowest bit precision may be the second and third layers.

[0158] Here, the first bit precision may be five-bit precision, and the second bit precision may be six-bit precision. Here, the processor 130 may quantize the second layer or the third layer in the neural network model 10 using six-bit precision, and quantize the remaining layers using the same bit precision as the first quantized neural network model 11. Therefore, the processor 130 may quantize the layers included in the neural network model 10 with [5 bits, 4 bits, 6 bits], and obtain the second quantized neural network model 13. Optionally, the processor 130 may quantize the layers included in the neural network model 10 with [5 bits, 6 bits, 4 bits], and obtain the second quantized neural network model 14. Optionally, the processor 130 may quantize the layers included in the neural network model 10 with [5 bits, 6 bits, 6 bits], and obtain the second quantized neural network model 15.

[0159] Optionally, the first quantized neural network model may be a neural network model 12, in which three layers are quantized in the order of [4 bits, 5 bits, 4 bits]. Here, the layers quantized using the lowest bit precision may be the first layer and the third layer.

[0160] Here, the first bit precision may be five-bit precision, and the second bit precision may be six-bit precision. Here, the processor 130 may quantize the first layer or the third layer in the neural network model 10 using the second bit precision (six-bit precision), and quantize the remaining layers using the same bit precision as the first quantized neural network model 11. Therefore, the processor 130 may quantize the layers included in the neural network model 10 with [6 bits, 5 bits, 4 bits], and obtain the second quantized neural network model 16. Optionally, the processor 130 may quantize the layers included in the neural network model 10 with [4 bits, 5 bits, 6 bits], and obtain the second quantized neural network model 17. Optionally, the processor 130 may quantize the layers included in the neural network model 10 with [6 bits, 5 bits, 6 bits], and obtain the second quantized neural network model 18. Subsequently, when multiple second quantized neural network models including a second quantized neural network model are acquired, the processor 130 may send the multiple second quantized neural network models including the second quantized neural network model to the target device 200 in operation S310, but this is merely an example, and the processor 130 may send the test data 20 together with the multiple second quantized neural network models including the second quantized neural network model to the target device 200.

[0161] Subsequently, the target device 200 may perform inference of each of a plurality of second quantized neural network models including the second quantized neural network model by using the test data 20. For example, the target device 200 may input the test data into each of the plurality of second quantized neural network models.

[0162] Here, the target device 200 may perform inference of each of the plurality of second quantized neural network models by using the test data, and obtain result data output by each of the plurality of second quantized neural network models. In addition, the target device 200 may perform inference of each of the plurality of second quantized neural network models by using the test data, and obtain profile information of the target device 200 related to each of the plurality of second quantized neural network models.

[0163] Subsequently, in operation S315 , the processor 130 may acquire, from the target device 200 , result data acquired from the quantized neural network model and profile information of the target device 200 related to the quantized neural network model.

[0164] Subsequently, the target device 200 may transmit the result data of each of the plurality of second quantized neural network models and the profile information of the target device 200 to the electronic device 100. For example, in operation S315, the processor 130 may acquire the result data acquired by performing inference of each of the plurality of second quantized neural network models by using the test data 20 and the profile information of the target device 200 from the target device 200.

[0165] Based on the acquired result data and the profile information of the target device 200, in operation S320, the processor 130 may identify a quantization error between the neural network model 10 and each of the plurality of second quantization neural network models.

[0166] Subsequently, the processor 130 may identify whether each of the plurality of second quantized neural network models meets the required performance condition. Specifically, the processor 130 may identify whether the profile information of the target device 200 for each of the plurality of second quantized neural network models meets the required performance condition.

[0167] Subsequently, in operation S325, the processor 130 may add a neural network model satisfying a required performance condition among the plurality of second quantized neural network models to the candidates 40 of available quantized neural network models.

[0168] Subsequently, in operation S330 , the processor 130 may identify whether the number of neural network models included in the candidates 40 of available quantization neural network models is less than a predetermined number k1.

[0169] Subsequently, if the number of neural network models included in the candidates 40 of available quantization neural network models is less than a predetermined number k1 (in operation S330-yes), in operation S340, the processor 130 may obtain a score for each of the remaining neural network models among the quantization neural network models in the plurality of second quantization neural network models except for the neural network model included in the candidates 40 of available quantization neural network models.

[0170] If the number of neural network models included in the candidates 40 of available quantization neural network models is less than a predetermined number k1 (in operation S330-yes), based on the profile information and the quantization error, the processor 130 may obtain a score for each of the neural network models among the multiple second quantization neural network models except the models included in the candidates 40 of available quantization neural network models.

[0171] Subsequently, based on the acquired scores, the processor 130 may identify a predetermined number k2 of neural network models among the neural network models other than the models included in the candidates 40 of available quantized neural network models among the plurality of second quantized neural network models.

[0172] Specifically, in operation S345, the processor 130 may identify a predetermined number k2 of neural network models in order of having higher scores among the neural network models excluding the models included in the candidates 40 of the available quantized neural network models among the plurality of second quantized neural network models.

[0173] Here, the processor may identify a neural network model that satisfies specific constraints among the neural network models in a plurality of second quantized neural network models other than the models included in the candidates 40 of available quantized neural network models, and identify a predetermined number k2 of neural network models among the neural network models that satisfy the specific constraints.

[0174] Subsequently, the electronic device 100 may repeatedly perform the operation of quantizing the neural network model and updating the available candidates 40 of the quantized neural network model by the method described above.

[0175] Therefore, until the number of models added to the available candidates 40 of the quantized neural network model becomes greater than or equal to the predetermined number k1, the processor 130 may repeatedly perform the operation of updating the available candidates 40 of the quantized neural network model.

[0176] Therefore, if the number of models included in the available candidates 40 of the quantized neural network model becomes greater than or equal to the predetermined number k1 (at operation S330-No), at operation S335, the processor 130 may identify a neural network model having the highest score among the neural network models included in the available candidates 40 of the quantized neural network model. Here, the method of acquiring a score for each of the neural network models included in the available candidates 40 of the quantized neural network model may be as described above.

[0177] Subsequently, when the neural network model with the highest score is identified among the candidates 40 of the available quantized neural network models, the processor 130 may transmit the identified neural network model to the target device 200. Here, the identified neural network model may be an available quantized neural network model. For example, the identified neural network model may be a neural network model that is quantized to be suitable for the target device 200.

[0178] Therefore, the electronic device 100 may acquire a neural network model quantized to be suitable for the target device 200 , and transmit the acquired quantized neural network model to the target device 200 .

[0179] In addition, the processor 130 may perform the operation of quantizing the neural network model 10 by various methods.

[0180] Specifically, according to an embodiment of the present disclosure, the (n+1)th bit precision may be a bit precision one level higher than the nth bit precision, but this is merely an example, and the (n+1)th bit precision may be a bit precision one level lower than the nth bit precision.

[0181] In the case where the (n+1)th bit precision is a bit precision one level higher than the nth bit precision, the quantization module 132 may quantize the neural network model 10 by the method described above (first method) at operations S305 and S350.

[0182] In addition, in the case where the (n+1)th bit precision is a bit precision one level lower than the nth bit precision, the quantization module 132 may quantize the neural network model 10 by the second method in operations S305 and S350.

[0183] Specifically, according to the second method, if n is 1 in operation S305, the processor 130 may quantize at least one layer among the plurality of layers included in the neural network model 10 using the first bit precision, and quantize the remaining layers except the at least one layer among the plurality of layers using the highest bit precision. Here, information about the highest bit precision may be stored in the memory 110. The first bit precision may be a bit precision one level lower than the highest bit precision.

[0184] In operation S305 , if n is 2 or greater, the processor 130 may quantize the neural network model 10 as in the operation in S350 to be described below.

[0185] In operation S350, the processor 130 may quantize at least one layer among the layers quantized using the highest bit precision in one model of the predetermined number k2 of neural network models using the (n+1)th bit precision. Subsequently, the processor 130 may quantize the layers other than the at least one layer among the plurality of layers using the same bit precision as the one model.

[0186] In addition, the processor 130 may quantize the neural network model 10 by using the first method or the second method, but this is merely an example, and the processor 130 may update the available candidates 40 for the quantized neural network model by using the first method and the second method together.

[0187] In addition, reference will be made to Figure 5 and Figure 6 A method for the electronic device 100 to acquire the profile information and the quantization error of the target device 200 in the aforementioned operations S305 , S310 , S315 , and S320 is described.

[0188] Figure 5 and Figure 6 is a diagram for illustrating a method for the electronic device 100 to acquire profile information and a quantization error of the target device 200 according to various embodiments of the present disclosure.

[0189] refer to Figure 5 and Figure 6 In operation S605, the quantization module 132 may obtain the neural network model 10, result data 11 of the neural network model 10 by using the test data 20, and information 12 about minimum and maximum values ​​of intermediate values ​​output by each layer included in the neural network model 10.

[0190] Subsequently, the quantization module 132 may store the obtained result data 11 of the neural network model 10 and information 12 about the minimum and maximum values ​​of intermediate values ​​output by each layer included in the neural network model 10 in the memory 110, but the present disclosure is not limited thereto.

[0191] Here, if the result data 11 and the information about the minimum and maximum values ​​12 are stored in the memory 110, the quantization module 132 may not perform the operation of acquiring the result data 11 and the information about the minimum and maximum values ​​12, and acquire the quantized neural network by using the information stored in the memory 110.

[0192] Subsequently, by using the information 12 about the minimum value and the maximum value, the quantization module 132 may quantize the intermediate values ​​output by each layer included in the neural network model 10 and quantize the weights included in each layer, and obtain a quantized neural network.

[0193] Specifically, the processor 130 may identify the quantization method and obtain the neural network model 13 quantized according to the identified quantization method. Here, the processor 130 may identify whether to quantize the neural network model 10 according to the identified quantization method based on whether information about the neural network model 13 quantized according to the identified quantization method is stored in the memory 110.

[0194] Specifically, in operation S610, the processor 130 may identify a quantization method for quantizing the neural network model 10. For example, the identified quantization method may be a method of quantizing the second layer of the neural network model 10 using a first bit precision and quantizing the remaining layers using a lowest bit precision.

[0195] Subsequently, at operation S615, the processor 130 may identify whether information about the neural network model 13 quantized according to the identified quantization method is stored in the memory 110. Here, the information about the quantized neural network model 13 may include a quantization error 16 between the quantized neural network model 13 and the neural network model 10, and profile information 15 of the target device 200 related to the quantized neural network model 13.

[0196] Subsequently, if the information about the neural network model 13 quantized according to the identified quantization method is stored in the memory 110 (at operation S615-Yes), then at operation S620, the processor 130 may obtain the profile information 15 related to the quantized neural network model 13 of the target device 200 and the quantization error 16 between the quantized neural network model 13 and the neural network model 10 by using the information about the quantized neural network model 13 stored in the memory 110. For example, if the information about the neural network model 13 quantized according to the identified quantization method is stored in the memory 110, the process of the electronic device 100 quantizing the neural network model 10 according to the identified quantization method and obtaining the profile information 15 and the quantization error 16 by using the quantized neural network model 13 may be omitted. Therefore, the resources and time incurred for the electronic device 100 to update the available candidates 40 of the quantized neural network model and obtain the neural network model quantized to be suitable for the target device 200 may be minimized.

[0197] If information about the neural network model 13 quantized according to the identified quantization method is not stored in the memory 110 (in operation S615-No), in operation S625, the processor 130 may quantize the neural network model 10 according to the identified quantization method and obtain the quantized neural network model 13.

[0198] Subsequently, the processor 130 may send the quantized neural network model 13 and the test data 20 to the target device 200 at operation S630, and acquire the result data 14 of the quantized neural network model 13 and the profile information 15 of the target device 200 from the target device 200 at operation S635. Subsequently, the processor 130 may store the acquired profile information of the target device 200 in the memory 110.

[0199] Then, in operation S640, the processor 130 may identify a quantization error 16 regarding the quantized neural network model 13 by using the acquired result data 14 and the result data 11 of the neural network model 10. Then, the processor 130 may store the acquired quantization error in the memory 110.

[0200] Figure 7 is a timing diagram for illustrating operations of an electronic device and a target device according to an embodiment of the present disclosure.

[0201] refer to Figure 7 In operation S710, the electronic device 100 may acquire a neural network model 10, test data 20 for the neural network model, and information 30 about required performance conditions.

[0202] Subsequently, in operation S720, the electronic device 100 may quantize the neural network model 10 using the n-th bit precision and obtain a first quantized neural network model.

[0203] Subsequently, the electronic device 100 may transmit the first quantized neural network model and the test data 20 to the target device 200 in operation S730 .

[0204] Therefore, the target device 200 may perform inference of each first quantized neural network model by using the test data 20. Therefore, in operation S740, the target device 200 may acquire result data acquired by performing inference of the first quantized neural network model and profile information of the target device 200 related to the first quantized neural network model.

[0205] Subsequently, the electronic device 100 may receive result data and profile information from the target device 200 in operation S750 .

[0206] When the result data and the profile information are received, the electronic device 100 may identify a quantization error of the first quantization neural network model by using the result data in operation S760.

[0207] Subsequently, if the first quantized neural network model satisfies the required performance condition, the electronic device 100 may add the first quantized neural network model to the candidates 40 of available quantized neural network models in operation S770.

[0208] Subsequently, the electronic device 100 and the target device 200 may repeatedly perform operations S720 to S760 until the number of models included in the candidates 40 of available quantized neural network models becomes greater than or equal to a predetermined number.

[0209] Subsequently, if the number of models included in the available candidates 40 of quantized neural network models becomes greater than or equal to a predetermined number, the electronic device 100 may transmit a model having a highest score among the models included in the available candidates 40 of quantized neural network models to the target device 200 in operation S780.

[0210] Figure 8 is a flowchart for illustrating a control method of an electronic device according to an embodiment of the present disclosure.

[0211] refer to Figure 8In operation S810, the electronic device 100 may acquire a neural network model 10, test data 20 for the neural network model, and information 30 about required performance conditions.

[0212] Subsequently, in operation S820, the electronic device 100 may quantize at least one of the multiple layers included in the neural network model 10 and obtain a first quantized neural network model. Here, the electronic device 100 may quantize at least one of the multiple layers included in the neural network model 10 using a first bit precision and obtain a first quantized neural network model.

[0213] Here, the electronic device 100 may quantize at least one layer among the multiple layers included in the neural network model 10, and obtain multiple first quantized neural network models including a first quantized neural network model.

[0214] Specifically, the electronic device 100 may change a layer to be quantized among a plurality of layers included in the neural network model 10, and obtain a plurality of first quantized neural network models including a first quantized neural network model.

[0215] Subsequently, the electronic device 100 may transmit the first quantized neural network model and the test data 20 to the target device 200 in operation S830 .

[0216] Subsequently, in operation S840, the electronic device 100 may receive, from the target device 200, result data acquired from the first quantized neural network model with the test data 20 as input, and profile information of the target device 200 related to the first quantized neural network model.

[0217] Here, the profile information of the target device 200 related to the first quantized neural network model may include at least one of the inference delay caused when the target device 200 performs inference of the first quantized neural network model, the memory usage of the target device 200, or the power consumption of the target device 200.

[0218] Thereafter, in operation S850, based on the result data and the profile information, if the first quantized neural network model satisfies the required performance condition, the electronic device 100 may add the first quantized neural network model to the candidates 40 of available quantized neural network models.

[0219] In addition, if the quantization error of the first quantization neural network model and profile information related to the first quantization neural network model of the target device are stored in the memory 110, the electronic device 100 can add the first quantization neural network model to the candidates 40 of available quantization neural network models by using the quantization error and profile information stored in the memory 110.

[0220] Subsequently, if the number of models included in the candidates 40 of available quantized neural network models is greater than or equal to a predetermined number, the electronic device 100 may obtain a score indicating suitability for the required performance condition for each of the neural network models included in the candidates 40 of available quantized neural network models. Subsequently, the electronic device 100 may identify a neural network model with the highest score among the neural network models included in the candidates 40 of available quantized neural network models. Subsequently, the electronic device 100 may send the identified neural network model with the highest score to the target device 200.

[0221] In addition, if the number of models included in the available candidates 40 for quantized neural network models is less than a predetermined number, and the first quantized neural network model does not meet the required performance conditions, the electronic device 100 may quantize the neural network model 10 using a second bit precision and obtain a second quantized neural network model.

[0222] Specifically, if the number of models included in the candidates 40 of available quantized neural network models is less than a predetermined number, the electronic device 100 may obtain a score indicating suitability for the required performance conditions for each of the models in the multiple first quantized neural network models other than the models included in the candidates 40 of available quantized neural network models.

[0223] Subsequently, the electronic device 100 may identify a predetermined number of neural network models in order of having higher scores among the neural network models other than the neural network models included in the candidates 40 of available quantized neural network models among the plurality of first quantized neural network models.

[0224] Here, the electronic device 100 may identify a neural network model that satisfies the restriction condition among the neural network models other than the models included in the candidates 40 of the available quantized neural network models among the plurality of first quantized neural network models. Subsequently, the electronic device 100 may acquire a score indicating suitability for the required performance condition for each of the neural network models that satisfy the restriction condition. Subsequently, the electronic device 100 may identify a predetermined number of neural network models in order of having higher scores among the neural network models that satisfy the restriction condition.

[0225] In addition, the electronic device 100 may identify the predetermined number based on at least one of the number of layers included in the neural network model 10 or information about the performance of the electronic device 100 .

[0226] Subsequently, based on the information about a predetermined number of neural network models, the electronic device 100 may quantize the neural network model 10 using a second bit precision and obtain a plurality of second quantized neural network models including a second quantized neural network model.

[0227] Specifically, based on information about the bit precision used to quantize a predetermined number of neural network models, the electronic device 100 can quantize at least one layer of the multiple layers included in the neural network model 10 using a second bit precision, and obtain multiple second quantized neural network models including a second quantized neural network model.

[0228] Subsequently, the electronic device 100 may send the second quantized neural network model to the target device 200. Based on the result data obtained from the second quantized neural network model and the profile information of the target device 200 related to the second quantized neural network model, if the second quantized neural network model meets the required performance conditions, the electronic device 100 may add the second quantized neural network model to the candidates 40 of the available quantized neural network models.

[0229] In addition, the term "component" or "module" used in the present disclosure may include a unit implemented as hardware, software or firmware, and may be used interchangeably with terms such as logic, logic block, component or circuit. In addition, a "component" or "module" may be a component that is configured to perform one or more functions, or its smallest unit, or a part thereof. For example, a module may be configured as an application specific integrated circuit (ASIC).

[0230] Various embodiments of the present disclosure may be implemented as software including instructions stored in a machine-readable storage medium, wherein the instructions can be read by a machine (e.g., a computer). A machine refers to a device that calls instructions stored in a storage medium and can operate according to the called instructions, and the device may include an electronic device 100 according to the aforementioned embodiment. In the case where instructions are executed by a processor, the processor may perform functions corresponding to the instructions by itself or by using other components under its control. Instructions may include code generated or executed by a compiler or interpreter. A machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory" only means that the storage medium does not include a signal and is tangible, without indicating whether the data is stored in the storage medium semi-permanently or temporarily.

[0231] In addition, according to one or more embodiments of the present disclosure, the methods according to the various embodiments described in the present disclosure may be provided in the case of being included in a computer program product. A computer program product refers to a product, and it can be traded between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or can be downloaded through an application store (e.g., PlayStore). TM) Online distribution. In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored in a storage medium (such as a memory of a manufacturer's server, an application store's server, and a relay server), or may be temporarily generated.

[0232] In addition, each component (e.g., module or program) according to various embodiments may include a single object or multiple objects. In addition, in the corresponding subcomponents mentioned above, some subcomponents may be omitted, or other subcomponents may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated as objects, and the functions performed by each component before integration are performed identically or in a similar manner. In addition, the operations performed by the modules, programs or other components according to various embodiments may be performed sequentially, in parallel, repeatedly or heuristically. Optionally, at least some operations may be performed or omitted in different orders, or other operations may be added.

[0233] While the present disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.

Claims

1. An electronic device, include: Communication interface; Memory; as well as at least one processor, Wherein, the at least one processor is configured to: obtaining a neural network model, test data for the neural network model, and information about desired performance conditions, quantizing at least one layer among the multiple layers included in the neural network model, and obtaining a first quantized neural network model, controlling the communication interface to send the first quantized neural network model and the test data to a target device, receiving, from the target device through the communication interface, result data obtained from the first quantitative neural network model using the test data as input, and configuration file information of the target device related to the first quantitative neural network model, and Based on the result data and the profile information, and based on the first quantized neural network model satisfying the required performance conditions, the first quantized neural network model is added to the available quantized neural network model candidates.

2. The electronic device according to claim 1, in, The at least one processor is further configured to: quantizing the at least one layer among the plurality of layers included in the neural network model using a first bit precision, and obtaining the first quantized neural network model, and Based on the fact that the number of neural network models included in the quantized neural network model candidates is less than a predetermined number and the first quantized neural network model does not meet the required performance conditions, the neural network model is quantized using a second bit precision and a second quantized neural network model is obtained.

3. The electronic device according to claim 2, in, The at least one processor is further configured to: sending the second quantized neural network model to the target device, receiving, from the target device, result data acquired from the second quantized neural network model and configuration file information of the target device related to the second quantized neural network model, and Based on the result data obtained from the second quantitative neural network model and the profile information of the target device related to the second quantitative neural network model, and based on the second quantitative neural network model satisfying the required performance conditions, the second quantitative neural network model is added to the available quantitative neural network model candidates.

4. The electronic device according to claim 1, in, The at least one processor is further configured to: changing a layer to be quantized among the plurality of layers included in the neural network model, and acquiring a plurality of first quantized neural network models including the first quantized neural network model, and Based on the fact that the number of neural network models included in the available quantized neural network model candidates is less than a predetermined number, a score indicating suitability for required performance conditions is obtained for each model of the multiple first quantized neural network models except for the models included in the available quantized neural network model candidates.

5. The electronic device according to claim 4, in, The at least one processor is further configured to: The predetermined number of neural network models are identified in order of higher scores among models other than the models included in the available quantized neural network model candidates.

6. The electronic device according to claim 5, in, The at least one processor is further configured to: The neural network model is quantized using a second bit precision based on a quantization method for quantizing the predetermined number of neural network models, and a second quantized neural network model is obtained.

7. The electronic device according to claim 4, in, The predetermined number is determined by at least one of the number of layers included in the neural network model or the performance of the electronic device.

8. The electronic device according to claim 4, in, The at least one processor is further configured to: identifying a neural network model satisfying a constraint condition among models other than the models included in the available quantized neural network model candidates, obtaining a score indicating suitability for the required performance condition for each of the neural network models satisfying the constraint condition, and The predetermined number of neural network models are identified in order of higher scores among the neural network models satisfying the restriction condition.

9. The electronic device according to claim 1, in, The at least one processor is further configured to: based on the number of neural network models included in the available quantized neural network model candidates being greater than or equal to a predetermined number, obtaining a score indicating suitability for a required performance condition for each of the neural network models included in the available quantized neural network model candidates, identifying a neural network model having a highest score among the neural network models included in the quantized neural network model candidates, and The communication interface is controlled to send the identified neural network model to the target device.

10. The electronic device according to claim 1, in, The at least one processor is further configured to: Based on the quantization error about the first quantization neural network model and the profile information related to the first quantization neural network model of the target device being stored in the memory, the first quantization neural network model is added to the available quantization neural network model candidates by using the quantization error and the profile information stored in the memory.

11. The electronic device according to claim 1, in, The configuration file information includes: At least one of an inference delay caused when the target device performs inference of the first quantized neural network model, a memory usage of the target device, or a power consumption of the target device.

12. A method for controlling an electronic device, the method include: obtaining a neural network model, test data for the neural network model, and information about desired performance conditions; quantizing at least one layer among a plurality of layers included in the neural network model, and obtaining a first quantized neural network model; Sending the first quantized neural network model and the test data to a target device; receiving, from the target device, result data obtained from the first quantitative neural network model using the test data as input, and configuration file information of the target device related to the first quantitative neural network model; as well as Based on the result data and the profile information, and based on the first quantized neural network model satisfying the required performance conditions, the first quantized neural network model is added to the available quantized neural network model candidates.

13. The method according to claim 12, in, Acquiring the first quantized neural network model includes: quantizing the at least one layer among the plurality of layers included in the neural network model using a first bit precision, and obtaining the first quantized neural network model, and Wherein, the method further comprises: Based on the fact that the number of neural network models included in the quantized neural network model candidates is less than a predetermined number and the first quantized neural network model does not meet the required performance conditions, the neural network model is quantized using a second bit precision and a second quantized neural network model is obtained.

14. The method according to claim 13, further comprising: include: sending the second quantized neural network model to the target device; receiving, from the target device, result data acquired from the second quantized neural network model and configuration file information of the target device related to the second quantized neural network model; as well as Based on the result data obtained from the second quantitative neural network model and the profile information of the target device related to the second quantitative neural network model, and based on the second quantitative neural network model satisfying the required performance conditions, the second quantitative neural network model is added to the available quantitative neural network model candidates.

15. The method according to claim 12, in, Acquiring the first quantized neural network model includes: changing a layer to be quantized among the plurality of layers included in the neural network model, and acquiring a plurality of first quantized neural network models including the first quantized neural network model, and Wherein, the method further comprises: Based on the fact that the number of neural network models included in the available quantized neural network model candidates is less than a predetermined number, a score indicating suitability for required performance conditions is obtained for each model of the multiple first quantized neural network models except for the models included in the available quantized neural network model candidates.