Electronic device and method for performing quantization on artificial intelligence models, and non-transitory computer-readable storage medium
Quantization of AI models into integer formats addresses memory and computation challenges, allowing efficient deployment on devices with limited resources by reducing model size and memory usage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-08-22
- Publication Date
- 2026-04-23
AI Technical Summary
Existing artificial intelligence models are not optimized for devices with limited memory capacity, leading to increased size, memory usage, and computation time, making them unsuitable for on-device applications.
Quantization of artificial intelligence models is performed to reduce model size, memory usage, and computation time by converting data from floating-point to integer formats, such as INT4, INT8, or INT16, using input, weight, and activation quantization parameters.
The quantization process significantly reduces the model size and memory usage, enabling efficient deployment of AI models on devices with limited resources while maintaining performance.
Smart Images

Figure KR2025012843_23042026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for performing quantization of artificial intelligence models
[0001] The following descriptions relate to an electronic device, a method, and a non-transient computer-readable storage medium for performing quantization of artificial intelligence models.
[0002] An electronic device may utilize an on-device artificial intelligence model. The on-device artificial intelligence model of the electronic device may be quantized in consideration of the device's limited memory capacity. By quantizing the inputs and weights of the on-device artificial intelligence model, the size, memory usage, and / or amount of computation of the on-device artificial intelligence model may be reduced.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0004] A server device is provided. The server device may include a communication circuit. The server device may include a memory that stores instructions and includes one or more storage media. The server device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the server device may cause at least one artificial intelligence model to identify by excluding an updatable artificial intelligence model from a plurality of artificial intelligence models of an ensemble model. When the instructions are executed individually or collectively by the at least one processor, the server device may cause quantization parameters to be determined based on performing quantization on a reference artificial intelligence model among the at least one artificial intelligence model. When the instructions are executed individually or collectively by the at least one processor, the server device may cause quantization on one or more artificial intelligence models excluding the reference artificial intelligence model among the plurality of artificial intelligence models based on the quantization parameters. When the above instructions are executed individually or collectively by the at least one processor, the server device may cause the server device to transmit a file for the ensemble model including the plurality of artificial intelligence models to an electronic device.
[0005] A method performed by a server device is provided. The method may include an operation of identifying at least one artificial intelligence model by excluding an updatable artificial intelligence model from a plurality of artificial intelligence models of an ensemble model. The method may include an operation of determining quantization parameters based on performing quantization on a reference artificial intelligence model among the at least one artificial intelligence model. The method may include an operation of performing quantization on one or more artificial intelligence models excluding the reference artificial intelligence model among the plurality of artificial intelligence models based on the quantization parameters. The method may include an operation of transmitting a file for the ensemble model including the plurality of artificial intelligence models to an electronic device.
[0006] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the server device to identify at least one artificial intelligence model by excluding an updatable artificial intelligence model from a plurality of artificial intelligence models of an ensemble model when executed individually or collectively by at least one processor of the server device. The one or more programs may include instructions that cause the server device to determine quantization parameters based on performing quantization on a reference artificial intelligence model among the at least one artificial intelligence model when executed individually or collectively by at least one processor of the server device. The one or more programs may include instructions that cause the server device to perform quantization on one or more artificial intelligence models excluding the reference artificial intelligence model among the plurality of artificial intelligence models based on the quantization parameters when executed individually or collectively by at least one processor of the server device. The above one or more programs may include instructions that cause the server device to transmit a file for the ensemble model, which includes the plurality of artificial intelligence models, to an electronic device when executed individually or collectively by at least one processor of the server device.
[0007] An electronic device is provided. The electronic device may include a communication circuit. The electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause a first output according to a first input by using a first artificial intelligence model among a plurality of artificial intelligence models stored in the memory. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause a second output according to a second input by using a second artificial intelligence model among the plurality of artificial intelligence models. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the output quantization parameters of the first artificial intelligence model and the output quantization parameters of the second artificial intelligence model to refrain from reading from the memory in order to perform de-quantization for the first output and the second output. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the input quantization parameters for the third artificial intelligence model among the plurality of artificial intelligence models to refrain from reading from the memory. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the third artificial intelligence model among the plurality of artificial intelligence models to generate an output based on the first output and the second output.
[0008] A method performed by an electronic device is provided. The method may include an operation of generating a first output according to a first input using a first artificial intelligence model among a plurality of artificial intelligence models stored in the memory. The method may include an operation of generating a second output according to a second input using a second artificial intelligence model among the plurality of artificial intelligence models. The method may include an operation of refraining from reading output quantization parameters of the first artificial intelligence model and output quantization parameters of the second artificial intelligence model from the memory in order to perform de-quantization of the first output and the second output. The method may include an operation of refraining from reading input quantization parameters for a third artificial intelligence model among the plurality of artificial intelligence models from the memory. The method may include an operation of generating an output based on the first output and the second output using a third artificial intelligence model among the plurality of artificial intelligence models.
[0009] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to generate a first output according to a first input by using a first artificial intelligence model among a plurality of artificial intelligence models stored in the memory when executed individually or collectively by at least one processor of the electronic device. The one or more programs may include instructions that cause the electronic device to generate a second output according to a second input by using a second artificial intelligence model among the plurality of artificial intelligence models when executed individually or collectively by at least one processor of the electronic device. The above one or more programs may include instructions that cause the electronic device to refrain from reading the output quantization parameters of the first artificial intelligence model and the output quantization parameters of the second artificial intelligence model from the memory in order to perform de-quantization of the first output and the second output when executed individually or collectively by at least one processor of the electronic device. The above one or more programs may include instructions that cause the electronic device to refrain from reading the input quantization parameters of the third artificial intelligence model among the plurality of artificial intelligence models from the memory when executed individually or collectively by at least one processor of the electronic device. The above one or more programs may include instructions that cause the electronic device to generate an output based on the first output and the second output using the third artificial intelligence model among the plurality of artificial intelligence models when executed individually or collectively by at least one processor of the electronic device.
[0010] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0011] Figure 1 is a block diagram of an electronic device in a network environment.
[0012] Figure 2 illustrates the components of an electronic device and a server.
[0013] FIGS. 3A and FIGS. 3B illustrate operations performed by an artificial intelligence model.
[0014] FIGS. 4a and FIGS. 4b illustrate a model structure in which multiple artificial intelligence models are connected.
[0015] Figure 5 is a flowchart illustrating the operations of a server device for obtaining a list of artificial intelligence models for determining quantization parameters.
[0016] FIG. 6 is a flowchart illustrating the operations of a server device for performing quantization on an ensemble model including multiple artificial intelligence models.
[0017] FIG. 7 is a flowchart illustrating the operations of a server device for performing quantization on an ensemble model including multiple artificial intelligence models.
[0018] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of other embodiments. A singular expression may include a plural expression unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art described in this disclosure. Terms used in this disclosure that are defined in a general dictionary may be interpreted as having the same or similar meaning as they have in the context of the relevant technology, and are not to be interpreted in an ideal or overly formal sense unless explicitly defined in this disclosure. In some cases, even terms defined in this disclosure are not to be interpreted to exclude the embodiments of this disclosure.
[0019] In the various embodiments of the present disclosure described below, a hardware-based approach is described as an example. However, since the various embodiments of the present disclosure include techniques using both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.
[0020] Additionally, in this disclosure, expressions of "greater than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled; however, this is merely for the purpose of expressing an example and does not exclude descriptions of "greater than" or "less than." Conditions described as "greater than" may be replaced with "greater than," conditions described as "less than" may be replaced with "less than," and conditions described as "greater than and less than" may be replaced with "greater than and less than." Furthermore, "A" to "B" below refer to at least one of elements from A (including A) to B (including B). Below, "C" and / or "D" refers to including at least one of "C" or "D," i.e., {"C", "D", "C" and "D"}.
[0021] Figure 1 is a block diagram of an electronic device in a network environment.
[0022] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0023] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0024] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0025] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, software (e.g., program (140)) and input or output data for related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0026] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0027] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0028] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0029] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0030] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0031] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0032] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0033] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0034] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0035] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0036] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0037] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0038] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0039] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0040] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0041] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0042] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0043] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0044] Figure 2 illustrates the components of an electronic device and a server.
[0045] Referring to FIG. 2, the system may include an electronic device (101) and a server device (210) to support an artificial intelligence model. In the following, the specific details regarding the processor (211), memory (212), and communication circuit (213) of the server device (210) may be substantially the same as the details regarding the processor (201), memory (202), and communication circuit (203) of the electronic device (101).
[0046] In FIG. 2, the electronic device (101) may include a processor (201), memory (202), communication circuit (203), and / or an artificial intelligence module (204). For example, the processor (201), memory (202), communication circuit (203), and artificial intelligence module (204) may be electronically and / or operably coupled with each other by a communication bus. Operatably coupled hardware components may mean that a direct or indirect connection between hardware components is established wired or wirelessly so that a second hardware component (e.g., memory (202), communication circuit (203), and / or artificial intelligence module (204)) is controlled by a first hardware component (e.g., processor (201)). Although the hardware components of the electronic device (101) illustrated in FIG. 2 are illustrated in different blocks, the present disclosure is not limited thereto. For example, some of the hardware components shown in FIG. 2 (e.g., a processor (201), memory (202), communication circuit (203), and / or at least some of the artificial intelligence module (204)) may be included in a single integrated circuit such as a system on chip (SoC) or a system in package (SIP). The type and number of hardware components included in the electronic device (101) are not limited to those shown in FIG. 2. For example, the electronic device (101) may include only some of the hardware components shown in FIG. 2.
[0047] In one embodiment, the electronic device (101) may include a processor (201). The processor (201) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), and a field programmable gate array (FPGA). As an example, the hardware component for processing data may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing unit (DSP), a microcontroller (MCU), and / or a neural processing unit (NPU). The number of processors (201) may be one or more. For example, the processor (201) may have the structure of a multi-core processor, such as a dual core, a quad core, or a hexa core. The processor (201) of FIG. 2 can have the same content as the processor (120) of FIG. 1 applied substantially.
[0048] In one embodiment, the processor (201) may include various processing circuits and / or a plurality of processors. For example, the term “processor” as used herein, including in the claims, may include various processing circuits including at least one processor, and one or more of the at least one processor may be configured to perform the various functions described below in a distributed manner, individually and / or collectively. As used below, where “processor,” “at least one processor,” and “one or more processors” are described as being configured to perform various functions, these terms encompass, for example, but not limited to, situations where one processor performs some of the cited functions and other processor(s) perform other parts of the cited functions, and also situations where one processor can perform all of the cited functions. Additionally, the at least one processor may include a combination of processors that perform the enumerated / disclosed various functions, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform the various functions.
[0049] In one embodiment, the electronic device (101) may include a memory (202). The memory (202) may include a hardware component for storing data and / or instructions that are input to or output from the processor (201). For example, the memory (202) may include a volatile memory such as random-access memory (RAM) and / or a non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, a hard disk, a compact disk, and an embedded multimedia card (eMMC).
[0050] In one embodiment, one or more instructions (or commands) representing operations and / or operations performed by the processor (201) of the electronic device (101) may be stored within the memory (202) of the electronic device (101). A set of one or more instructions may be referred to as a program, firmware, operating system, process, routine, sub-routine, and / or application. Hereinafter, being installed within the electronic device (101) may mean that one or more instructions provided in the form of an application are stored within the memory (202), and that one or more applications are stored in an executable format by the processor (201) of the electronic device (101). The specific details regarding the memory (202) of FIG. 2 may be substantially the same as the details regarding the memory (130) of FIG. 1.
[0051] In one embodiment, the electronic device (101) may include a communication circuit (203). The communication circuit (203) may include a circuit for supporting the transmission and / or reception of electrical signals between the electronic device (101) and an external electronic device different from the electronic device (101) (e.g., a server (210)). The communication circuit (203) may include at least one of a modem, an antenna, and an O / E (optic / electronic) converter. The communication circuit (203) may support the transmission and / or reception of electrical signals based on various types of communication means such as Ethernet, Bluetooth, BLE (Bluetooth Low Energy), ZigBee, LTE (Long Term Evolution), and 5G NR (New Radio). The specific details regarding the communication circuit (203) of FIG. 2 may be substantially the same as the details regarding the communication module (109) and / or antenna module (197) of FIG. 1.
[0052] In one embodiment, the electronic device (101) may include an artificial intelligence module (204). The artificial intelligence module (204) may be a unit (function, code, separate device, circuit, or set of instructions) for performing functions. For example, the artificial intelligence module (204) may be a unit (function, code, separate device, circuit, or set of instructions) for a vision encoder (VE), a projector, a large language model (LLM), an embedding, and / or a large multimodal model (LMM). Hereinafter, the artificial intelligence module (207) may be referred to as an artificial intelligence model or other terms having an equivalent technical / functional meaning. The specific details regarding the artificial intelligence module (207) of FIG. 2 may be substantially identical to the details regarding the artificial intelligence model (300) of FIG. 3a and FIG. 3b.
[0053] FIGS. 3A and 3B illustrate operations performed by an artificial intelligence model. FIGS. 3A and 3B explain the need for quantization for an on-device artificial intelligence model stored in an electronic device (101). For example, the artificial intelligence model (300) may be an on-device artificial intelligence model stored in an electronic device (101).
[0054] Referring to FIG. 3a, the artificial intelligence model (300) can perform a linear operation (or linear transformation) (301), a Conv2D (convolution two-dimensional) operation (302), a Conv2D operation (303), and a Matmul (matrix multiplication) operation (304). For example, the artificial intelligence model (300) can perform a linear operation (301) on an input. The artificial intelligence model (300) can perform a Conv2D operation (302) and a Conv2D operation (303) on the result (or output tensor) of the linear operation (301). The artificial intelligence model (300) can concatenate the result (or output tensor) of the Conv2D operation (302) and the result (or output tensor) of the Conv2D operation (303). The artificial intelligence model (300) can concatenate the concatenated By performing Matmul operations (304) on the results, output can be generated (or inferred). In FIG. 3a, an artificial intelligence model (300) that performs linear operations (301), Conv2D operations (302), Conv2D operations (303), and Matmul operations (304) is described, but this is merely an example and the present disclosure is not limited thereto. For example, the artificial intelligence model (300) may be constructed based on the operations described above and / or other operations.
[0055] In one embodiment, the artificial intelligence model (300) may be an FP32 (floating point) model. An FP32 model may use a 32-bit floating point format to represent data. In one example, the data of the FP32 model may be represented based on 1 sign bit, 8 exponent bits, and 23 fraction bits. Since the data of the FP32 model is represented based on 32 bits, the actual value to be represented by the data can be represented relatively accurately. The data of the FP32 model may have a size of 4 bytes. The FP32 model may perform operations (e.g., linear operation (301), Conv2D operation (302), and Conv2D operation (303)) using weights learned through a training process. However, since each data (or actual value) in the FP32 model is represented based on 32 bits, the model size and memory usage may be relatively large. Because the model size and memory usage are relatively large, the time required for the artificial intelligence model (300) to perform operations may be relatively long. Therefore, the FP32 model may not be suitable as an on-device artificial intelligence model used in the electronic device (101).
[0056] According to one embodiment, quantization may be used to implement an on-device artificial intelligence model of an electronic device (101) in consideration of model size, memory usage, and / or computation time. In one example, an INT4 model may use 4 bits to represent data. The INT4 model may have 12.5% of the model size and / or memory usage compared to an FP32 model. In one example, an INT8 model may use 8 bits to represent data. The INT8 model may have 25% of the model size and / or memory usage compared to an FP32 model. In one example, an INT16 model may use 16 bits to represent data. The INT16 model may have 50% of the model size and / or memory usage compared to an FP32 model.
[0057] Referring to FIG. 3b, in one embodiment, the artificial intelligence model (300) may be a quantization model. For example, the artificial intelligence model (300) may perform input quantization (311), weight quantization (312), and activation quantization (313) for linear operation (301).
[0058] For example, the artificial intelligence model (300) can perform input quantization (311). Input quantization (311) can be performed based on input quantization parameters (e.g., scale and offset). A dataset for the input quantization parameters can be stored in memory (202). For example, the artificial intelligence model (300) can obtain input quantization parameters from memory (202). Based on the input quantization parameters, the artificial intelligence model (300) can convert an input having an FP32 data type into an input having an integer data type (e.g., INT(integer)4, INT8, or INT16).
[0059] For example, the artificial intelligence model (300) can perform weight quantization (312). Weight quantization (312) can be performed based on weights and weight quantization parameters (e.g., scale and offset). A data set for weights and / or weight quantization parameters can be stored in memory (202). For example, the artificial intelligence model (300) can obtain weights and / or weight quantization parameters from memory (202). The artificial intelligence model (300) can convert weights into quantized weights based on the weight quantization parameters. The artificial intelligence model (300) can perform linear operations (301) based on inputs converted into integer data formats based on input quantization (311) and quantized weights.
[0060] For example, the artificial intelligence model (300) can perform activation quantization (313). Activation quantization (313) can be performed based on activation quantization parameters (e.g., scale and offset). A data set for the activation quantization parameters can be stored in memory (202). For example, the artificial intelligence model (300) can obtain the activation quantization parameters from memory (202). Based on the activation quantization parameters, the artificial intelligence model (300) can convert the result of a linear operation (301) into an integer data format to be used as an input for Conv2D operations (e.g., Conv2D operation (302) and Conv2D operation (303)). In one example, the artificial intelligence model (300) can convert the result of a linear operation (301) having a first integer data format (e.g., INT4) into a result having a second integer data format (e.g., INT8). The artificial intelligence model (300) can perform Conv2D operations (302) and Conv2D operations (303) based on the result of a linear operation (301) in which activation quantization (313) is performed. For example, the result of the Conv2D operation (302) and Conv2D operation (303) can be used as input for a Matmul operation (304).
[0061] Referring to FIG. 3b, in one embodiment, an artificial intelligence model (300) may perform input quantization (321) and output quantization (322) for a Matmul operation (304). For example, the input quantization (321) may correspond to a quantized operation result based on activation quantization parameters (e.g., scale and offset) of a Conv2D operation (302) and a quantized operation result based on quantization parameters (e.g., scale and offset) of a Conv2D operation (303). In one example, the activation quantization parameters of the Conv2D operation (302) and the quantization parameters of the Conv2D operation (303) may be the same or different.
[0062] For example, the artificial intelligence model (300) can perform output quantization (322). Output quantization (322) can be performed based on output quantization parameters (e.g., scale and offset). A data set for the output quantization parameters can be stored in memory (202). For example, the artificial intelligence model (300) can obtain output quantization parameters from memory (202). Based on the output quantization parameters, the artificial intelligence model (300) can convert the result of a Matmul operation (304) having an integer data format into an output having an FP32 data format.
[0063] FIGS. 4A and 4B illustrate a model structure in which multiple artificial intelligence models are connected. For example, the multiple artificial intelligence models may be referred to as an ensemble model, a super-model, or other terms having an equivalent technical or functional meaning. FIGS. 4A and 4B describe a model structure for minimizing the intervention of the processor (201) of the electronic device (101) (or the CPU (central processing unit) of the processor (201)) in an ensemble model.
[0064] Referring to FIG. 4a, the ensemble model may include a first artificial intelligence model (410), a second artificial intelligence model (420), and a third artificial intelligence model (430). For example, in the model structure illustrated in FIG. 4a, the output of the first artificial intelligence model (410) and the output of the second artificial intelligence model (420) may be used as inputs to the third artificial intelligence model (430). In one example, the third artificial intelligence model (430) that generates (or infers) the output may be referred to as the main model, main artificial intelligence model, or other terms having an equivalent technical / functional meaning. In one example, the first artificial intelligence model (410) or the second artificial intelligence model (420) that obtains the input may be referred to as the sub model, sub artificial intelligence model, or other terms having an equivalent technical / functional meaning. However, this is merely an example and the present disclosure is not limited thereto.
[0065] In the model structure illustrated in FIG. 4a, the first artificial intelligence model may be a quantized model based on first quantization parameters (e.g., a first scale and a first offset). The first quantization parameters may be output quantization parameters of the first artificial intelligence model. The second artificial intelligence model may be a quantized model based on second quantization parameters (e.g., a second scale and a second offset). The second quantization parameters may be output quantization parameters of the second artificial intelligence model. The third artificial intelligence model may be a quantized model based on third quantization parameters (e.g., a third scale and a third offset). The second quantization parameters may be input quantization parameters of the third artificial intelligence model. For example, the first quantization parameters, the second quantization parameters, and the third quantization parameters may be different from each other.
[0066] In one example, de-quantization may be performed according to [Equation 1] below. However, this is merely an example and the present disclosure is not limited thereto. For example, de-quantization may be performed in a manner other than [Equation 1] below.
[0067]
[0068] x represents the value (e.g., the actual value) expressed based on the FP32 data format. Scale represents the scale among the quantization parameters. Q represents the quantized value. Offset represents the offset among the quantization parameters.
[0069] For example, the first artificial intelligence model (410) can generate (or infer) a first output based on a first input. The first output may be quantized values based on the first quantization parameters of the first artificial intelligence model (410). The first quantization parameters of the first artificial intelligence model (410) may be different from the third quantization parameters of the third artificial intelligence model (430). According to [Equation 1], the actual value (x) (e.g., FP32 value) identified based on the first quantization parameters for the first output (Q) and the actual value (x) identified based on the third quantization parameters for the first output (Q) may be different. Since the actual value may be interpreted differently in the first artificial intelligence model (410) and the third artificial intelligence model (430), de-quantization (415) and quantization (416) need to be performed in order to use the first output as an input to the third artificial intelligence model (430). De-quantization (415) and quantization (416) may be performed by the processor (201) of the electronic device (101) (or the CPU (central processing unit) of the processor (201)). For example, the processor (201) may perform de-quantization on the first output based on the first quantization parameters. By performing de-quantization on the first output, the processor (201) may obtain values in the FP32 data format for the first output. The processor (201) can obtain values of the data format of the third artificial intelligence model (430) by quantizing values of the FP32 data format based on the third quantization parameters.
[0070] For example, the second artificial intelligence model (420) can generate (or infer) a second output based on the second input. The second output may be quantized values based on the second quantization parameters of the second artificial intelligence model (420). The second quantization parameters of the second artificial intelligence model (420) may be different from the third quantization parameters of the third artificial intelligence model (430). According to [Equation 1], the actual value identified based on the second quantization parameters for the second output and the actual value identified based on the third quantization parameters for the second output may be different. Since the actual value may be interpreted differently in the second artificial intelligence model (420) and the third artificial intelligence model (430), non-quantization (425) and quantization (426) need to be performed to provide the second output as an input to the third artificial intelligence model (430). De-quantization (425) and quantization (426) can be performed by the processor (201) of the electronic device (101) (or the CPU of the processor (201)). For example, the processor (201) can perform de-quantization for the second output based on second quantization parameters. By performing de-quantization for the second output, the processor (201) can obtain values of the FP32 data format for the second output. The processor (201) can obtain values of the data format of the third artificial intelligence model (430) by quantizing the values of the FP32 data format based on third quantization parameters.
[0071] As described above, in an ensemble model comprising multiple artificial intelligence models, non-quantization (e.g., non-quantization (415), non-quantization (416)) and quantization (e.g., quantization (416), quantization (426)) need to be performed by the processor (201) of the electronic device (101) in order to match the values interpreted in each artificial intelligence model. Intervention by the processor (201) to perform non-quantization and quantization can cause latency in generating the output of the ensemble model by increasing the amount of computation of the electronic device (101). To eliminate such intervention by the processor (201), the output quantization parameters of the first artificial intelligence model (410) and the second artificial intelligence model (420) and the input quantization parameters of the third artificial intelligence model (430) need to be set to be identical.
[0072] Referring to FIG. 4b, the ensemble model may include a first artificial intelligence model (410), a second artificial intelligence model (420), and a third artificial intelligence model (430). In the model structure illustrated in FIG. 4b, the output of the first artificial intelligence model (410) and the output of the second artificial intelligence model (420) may be used as inputs to the third artificial intelligence model (430). In one example, the third artificial intelligence model (430) that generates (or infers) the output may be referred to as the main model. In one example, the first artificial intelligence model (410) or the second artificial intelligence model (420) that obtains the input may be referred to as the sub-model. However, this is merely an example and the present disclosure is not limited thereto.
[0073] In the model structure illustrated in FIG. 4b, the first artificial intelligence model (410) may be a quantized model based on quantization parameters (411). The quantization parameters (411) may be for output quantization of the first artificial intelligence model (410). The second artificial intelligence model (420) may be a quantized model based on quantization parameters (421). The quantization parameters (421) may be for output quantization of the second artificial intelligence model (420). The third artificial intelligence model (430) may be a quantized model based on quantization parameters (431). The quantization parameters (431) may be for input quantization of the third artificial intelligence model (430). For example, the quantization parameters (411), quantization parameters (421), and quantization parameters (431) may be identical to each other.
[0074] For example, the first artificial intelligence model (410) can generate (or infer) a first output based on the first input. The first output may be quantized values based on the quantization parameters (411) of the first artificial intelligence model (410). The quantization parameters (411) of the first artificial intelligence model (410) may be the same as the quantization parameters (431) of the third artificial intelligence model (430). According to [Equation 1], the actual value (e.g., FP32 value) identified based on the quantization parameters (411) for the first output and the actual value identified based on the quantization parameters (431) for the first output may be the same. Since the actual value can be interpreted identically in the first artificial intelligence model (410) and the third artificial intelligence model (430), the de-quantization (415) and quantization (416) described in FIG. 4a may not be performed in order to use the first output as an input to the third artificial intelligence model (430). Since the de-quantization (415) and quantization (416) are not performed, the intervention of the processor (201) may be eliminated. For example, the electronic device (101) may refrain from reading the quantization parameters (411) of the first artificial intelligence model (401) from memory (202) in order to perform de-quantization on the first output of the first artificial intelligence model (410). For example, the electronic device (101) may refrain from reading the quantization parameters (431) of the third artificial intelligence model (430) from memory (202).
[0075] For example, the second artificial intelligence model (420) can generate (or infer) a second output based on the second input. The second output may be quantized values based on the quantization parameters (421) of the second artificial intelligence model (420). The quantization parameters (421) of the second artificial intelligence model (420) may be the same as the quantization parameters (431) of the third artificial intelligence model (430). According to [Equation 1], the actual value identified based on the quantization parameters (411) for the second output and the actual value identified based on the quantization parameters (431) for the second output may be the same. Since the actual value can be interpreted identically in the first artificial intelligence model (410) and the third artificial intelligence model (430), the non-quantization (425) and quantization (426) described in FIG. 4a may not be performed in order to use the second output as an input to the third artificial intelligence model (430). Since the non-quantization (425) and quantization (426) are not performed, the intervention of the processor (201) may be eliminated. For example, the electronic device (101) may refrain from reading the quantization parameters (421) of the second artificial intelligence model (401) from memory (202) to perform non-quantization for the second output of the second artificial intelligence model (420). For example, the electronic device (101) may refrain from reading the quantization parameters (431) of the third artificial intelligence model (430) from memory (202).
[0076] As described above, when the quantization parameters (411), quantization parameters (421), and quantization parameters (431) are identical, the intervention of the processor (201) for performing non-quantization and quantization can be eliminated. Below, an electronic device (101) and a server device (210) for determining a reference artificial intelligence model for determining the quantization parameters (411), quantization parameters (421), and quantization parameters (431) are described.
[0077] FIG. 5 is a flowchart illustrating the operations of a server device for obtaining a list of artificial intelligence models for determining quantization parameters. The operations of FIG. 5 can be performed by the server device (210) of FIG. 2. For example, at least some of the operations can be controlled by the processor (211) of the server device (210). In the following, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed. For example, at least two operations may be performed in parallel.
[0078] Referring to FIG. 5, in operation 501, a server device (210) according to one embodiment can identify whether an artificial intelligence model among a plurality of artificial intelligence models is an updatable artificial intelligence model.
[0079] In one embodiment, a plurality of artificial intelligence models may constitute an ensemble model. In one example, the ensemble model may be referred to as a super-model or another term having an equivalent technical / functional meaning. For example, the ensemble model may have a form in which the output of at least one artificial intelligence model among the plurality of artificial intelligence models is used as the input of one artificial intelligence model. In another example, the ensemble model may have a form in which the output of one artificial intelligence model among the plurality of artificial intelligence models is used as the input of at least one artificial intelligence model. In one example, the ensemble model may have the model structure illustrated in FIG. 4b. The plurality of artificial intelligence models of the ensemble model may include a first artificial intelligence model, a second artificial intelligence model, and a third artificial intelligence model. The output of the first artificial intelligence model and the output of the second artificial intelligence model may be used as the input of the third artificial intelligence model. The first artificial intelligence model or the second artificial intelligence model may be referred to as a sub-model, a sub-artificial intelligence model, or another term having an equivalent technical / functional meaning. The third artificial intelligence model may be referred to as the main model, main artificial intelligence model, or other terms having an equivalent technical or functional meaning. However, this is merely an example to explain the structure of the ensemble model, and the present disclosure is not limited thereto. For example, the plurality of artificial intelligence models may include two or fewer or four or more artificial intelligence models. For example, the plurality of artificial intelligence models may have a connection structure different from the structure described above.
[0080] In one embodiment, the quantization parameters described in FIG. 4b (e.g., scale and offset) may be obtained (or determined) based on quantization of a reference AI model among a plurality of AI models. For example, when an update is performed on the reference AI model, the quantization parameters may also be changed. When the quantization parameters of the reference AI model are changed, quantization may be performed to change the quantization parameters of the remaining AI models among the plurality of AI models. To prevent re-quantization from being performed to change the quantization parameters of the remaining AI models, the updateable AI model may not be used as the reference AI model for determining the quantization parameters. For example, the updateable AI model may be removed from the list associated with the reference AI model. The updateable AI model may cause changes in the quantization parameters of the remaining AI models when the quantization parameters of the corresponding AI model are changed. To prevent re-quantization from being performed on the remaining models, the updateable AI model may be removed from the list. The updateable model may perform quantization based on QAT (quantization aware training) or PTQ (post training quantization) using quantization parameters determined based on the reference AI model, and generate a frozen graph. For example, a non-updatable model may be used as the reference AI model for determining quantization parameters.In one example, an unupdatable model may be referred to as a fixed model, a fixed artificial intelligence model, or other terms having an equivalent technical or functional meaning. In one example, an updatable model or an unupdatable artificial intelligence model may be predefined by an enterprise (or engineer) among multiple artificial intelligence models. In one example, an unupdatable model may have fixed weights and a fixed architecture (e.g., number of layers, number of nodes, connection method).
[0081] In operation 502, a server device (210) according to one embodiment can identify the next artificial intelligence model among a plurality of artificial intelligence models. For example, the server device (210) can identify the next artificial intelligence model among a plurality of artificial intelligence models based on the identification that the artificial intelligence model is an updateable artificial intelligence model. For example, an updateable artificial intelligence model may not be used as a reference artificial intelligence model for determining quantization parameters. The server device (210) can perform operations according to FIG. 5 starting from operation 501 for the next artificial intelligence model.
[0082] In operation 503, a server device (210) according to one embodiment can identify whether an artificial intelligence model can perform quantization based on quantization-aware training (QAT). For example, the server device (210) can identify whether the artificial intelligence model can perform quantization based on QAT based on the identification that the artificial intelligence model is not an updatable artificial intelligence model.
[0083] For example, quantization methods may include QAT and PTQ (post-training quantization). For instance, QAT may be a method of acquiring quantization parameters by performing quantization and training together. Since QAT performs both quantization and training, the time required to perform quantization may be relatively long. For instance, PTQ may be a method of acquiring quantization parameters by performing quantization on an AI model that has completed training. Since PTQ performs only quantization, the time required to perform quantization according to PTQ may be relatively short. Therefore, AI models with relatively large model sizes (e.g., LLM (large language model)) may practically only perform PTQ. To prevent latency in acquiring quantization parameters, AI models with relatively large model sizes may not be used as reference AI models. For instance, an AI model capable of only PTQ (available, or supporting) may not be used as a reference AI model for determining quantization parameters. For example, a QAT-capable AI model can be used as a reference AI model for determining quantization parameters. In one example, a QAT-capable AI model can be identified based on the model size. A server device (210) can identify an AI model as a PTQ-only AI model based on the identification that the size of the AI model exceeds a threshold size. A server device (210) can identify an AI model as a QAT-capable AI model based on the identification that the size of the AI model is less than a threshold size. A threshold size can be predefined to identify whether the AI model is PTQ-only.
[0084] In operation 504, a server device (210) according to one embodiment may add a corresponding artificial intelligence model to a list. For example, the server device (210) may add a corresponding artificial intelligence model to a list based on the identification that the artificial intelligence model can perform quantization based on QAT. The list may include artificial intelligence models available to determine the quantization parameters described in FIG. 4b. The list may be sorted in ascending order according to model size. The model size may correspond to the number of operations of the corresponding artificial intelligence model and / or the size of the file of the corresponding artificial intelligence model. In another example, the server device (210) may identify the next artificial intelligence model based on the identification that the artificial intelligence model cannot perform quantization based on QAT. The server device (210) may perform operations according to FIG. 5 starting from operation 501 for the next artificial intelligence model.
[0085] FIG. 6 is a flowchart illustrating the operations of a server device for performing quantization on an ensemble model comprising multiple artificial intelligence models. The operations of FIG. 6 can be performed by the server device (210) of FIG. 2. For example, at least some of the operations can be controlled by the processor (211) of the server device (210). In the following, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed. For example, at least two operations may be performed in parallel. For example, the operations described in FIG. 6 may be performed subsequently to the operations described in FIG. 5.
[0086] Referring to FIG. 6, in operation 601, a server device (210) according to one embodiment can obtain quantization parameters based on an artificial intelligence model of a list.
[0087] In one embodiment, the server device (210) can identify an artificial intelligence model having a minimum model size among the artificial intelligence models included in the list. For example, the list may include artificial intelligence models available for determining quantization parameters. The list may be sorted in ascending order according to model size. The model size may correspond to the number of operations of the corresponding artificial intelligence model and / or the size of the file of the corresponding artificial intelligence model. For example, an artificial intelligence model having a minimum model size may be a criterion or reference artificial intelligence model for obtaining (or determining) quantization parameters.
[0088] In one embodiment, the server device (210) can perform quantization on an artificial intelligence model having a minimum model size. Quantization methods may include quantization-aware training (QAT) and post-training quantization (PTQ). For example, the server device (210) may select one quantization method from QAT or PTQ based on the model structure and / or complexity. For example, the server device (210) may identify a quantization method based on the model size. In one example, the server device (210) may perform quantization on the artificial intelligence model based on PTQ upon identifying that the model size exceeds a threshold size. By performing quantization on the artificial intelligence model based on PTQ, the time required to perform quantization may be reduced. In one example, the server device (210) may perform quantization on the artificial intelligence model based on QAT upon identifying that the model size is less than or equal to a threshold size. By performing quantization on the artificial intelligence model based on QAT, relatively accurate quantization parameters can be determined. A threshold size can be predefined to determine the quantization method. After performing quantization, the server device (210) can generate a frozen graph for the first artificial intelligence model.
[0089] In operation 602, a server device (210) according to one embodiment may perform quantization of the remaining artificial intelligence models based on quantization parameters. For example, the server device (210) may perform quantization of the remaining artificial intelligence models excluding the reference artificial intelligence model from a plurality of artificial intelligence models, and then generate frozen graphs for the remaining artificial intelligence models.
[0090] In one embodiment, a plurality of artificial intelligence models may constitute an ensemble model. In one example, the ensemble model may be referred to as a super-model or another term having an equivalent technical / functional meaning. The ensemble model may have a form in which the output of at least one artificial intelligence model among the plurality of artificial intelligence models is used as the input of another artificial intelligence model. In one example, the plurality of artificial intelligence models may include a first artificial intelligence model, a second artificial intelligence model, and a third artificial intelligence model. The output of the first artificial intelligence model and the output of the second artificial intelligence model may be used as the input of the third artificial intelligence model. In one example, in the ensemble model described above, the first artificial intelligence model may be a reference artificial intelligence model. The server device (210) may obtain (or determine) output quantization parameters by performing quantization on the first artificial intelligence model. The server device (210) may perform quantization on the second artificial intelligence model such that the output quantization parameters of the second artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT. The server device (210) may perform quantization for the third artificial intelligence model such that the input quantization parameters of the third artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT. However, this is merely an example to explain the structure of the ensemble model, and the present disclosure is not limited thereto. The structure of the ensemble model according to the present disclosure may have various forms in which the output of at least one artificial intelligence model among a plurality of artificial intelligence models is used as the input of another artificial intelligence model. For example, the output of the first artificial intelligence model may be used as the input of the second artificial intelligence model and the input of the third artificial intelligence model. In one example, the first artificial intelligence model in the above-described ensemble model may be a reference artificial intelligence model.The server device (210) can obtain (or determine) output quantization parameters by performing quantization on the first artificial intelligence model. The server device (210) can perform quantization on the second artificial intelligence model such that the input quantization parameters of the second artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT. The server device (210) can perform quantization on the third artificial intelligence model such that the input quantization parameters of the third artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT.
[0091] In operation 603, a server device (210) according to one embodiment may determine whether the accuracy of an ensemble model exceeds a threshold value. For example, the server device (210) may determine the accuracy of an ensemble model after generating frozen graphs for a plurality of artificial intelligence models. For example, the accuracy of an ensemble model may correspond to a quantized signal-to-noise ratio (QSNR) or a custom metric. For example, a threshold value for evaluating accuracy may be predefined.
[0092] In operation 604, a server device (210) according to one embodiment may transmit a file of an ensemble model to an electronic device (101). For example, the server device (210) may generate a file of an ensemble model including a plurality of artificial intelligence models based on a determination that the accuracy of the ensemble model exceeds a threshold value. The server device (210) may transmit the generated file to the electronic device (101).
[0093] In operation 605, a server device (210) according to one embodiment can identify the next artificial intelligence model. For example, the server device (210) can identify the next artificial intelligence model based on a determination that the accuracy of the ensemble model is below a threshold value. For example, the next artificial intelligence model may be an artificial intelligence model having the second lowest model size among the artificial intelligence models included in the list. The server device (210) can perform operations according to FIG. 6 on the next artificial intelligence model starting from operation 601.
[0094] FIG. 7 is a flowchart illustrating the operations of a server device for performing quantization on an ensemble model comprising multiple artificial intelligence models. The operations of FIG. 7 can be performed by the server device (210) of FIG. 2. For example, at least some of the operations can be controlled by the processor (211) of the server device (210). In the following, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed. For example, at least two operations may be performed in parallel.
[0095] Referring to FIG. 7, in operation 701, a server device (210) according to one embodiment can identify at least one artificial intelligence model among a plurality of artificial intelligence models.
[0096] In one embodiment, a plurality of artificial intelligence models may constitute an ensemble model. In one example, the ensemble model may be referred to as a super-model or another term having an equivalent technical / functional meaning. For example, the ensemble model may have a form in which the output of at least one of the plurality of artificial intelligence models is used as the input of one artificial intelligence model. In another example, the ensemble model may have a form in which the output of one of the plurality of artificial intelligence models is used as the input of at least one artificial intelligence model. In one example, the plurality of artificial intelligence models may have the model structure illustrated in FIG. 4b. The plurality of artificial intelligence models may include a first artificial intelligence model (e.g., the first artificial intelligence model (410) of FIG. 4b), a second artificial intelligence model (e.g., the second artificial intelligence model (420) of FIG. 4b), and a third artificial intelligence model (e.g., the third artificial intelligence model (430) of FIG. 4b). The output of the first artificial intelligence model and the output of the second artificial intelligence model may be used as the input of the third artificial intelligence model. The first artificial intelligence model or the second artificial intelligence model may be referred to as a sub-model, sub-artificial intelligence model, or other terms having an equivalent technical or functional meaning. The third artificial intelligence model may be referred to as a main model, main artificial intelligence model, or other terms having an equivalent technical or functional meaning. However, the structure described above is merely an example, and the present disclosure is not limited thereto. For example, the plurality of artificial intelligence models may include two or fewer or four or more artificial intelligence models. For example, the structure of the plurality of artificial intelligence models may differ from the connection structure described above.
[0097] In one embodiment, the electronic device (101) can identify at least one artificial intelligence model by excluding an updatable artificial intelligence model from a plurality of artificial intelligence models. For example, quantization parameters (e.g., scale and offset) may be determined (or obtained) based on quantization for a criterion (or reference) artificial intelligence model. When an update is performed on the criterion artificial intelligence model, the quantization parameters may also be changed. When the quantization parameters of the criterion artificial intelligence model are changed, quantization needs to be performed to change the quantization parameters for the remaining artificial intelligence models among the plurality of artificial intelligence models, excluding the criterion artificial intelligence model. Thus, to prevent re-quantization to change the quantization parameters for the remaining artificial intelligence models, the updatable artificial intelligence model may not be used as the criterion artificial intelligence model for determining the quantization parameters. For example, a non-updatable artificial intelligence model may be used as the criterion artificial intelligence model for determining the quantization parameters. In one example, an unupdatable AI model may be referred to as a fixed model, a fixed AI model, or other terms having an equivalent technical or functional meaning. In one example, an unupdatable AI model may be predefined by an enterprise (or developer) among multiple AI models. In one example, an unupdatable model may have fixed weights and a fixed architecture (e.g., number of layers, number of nodes, connection method).
[0098] In one embodiment, the electronic device (101) can identify at least one artificial intelligence model by identifying an artificial intelligence model capable of quantization-aware training (QAT) among a plurality of artificial intelligence models. For example, the quantization method may include QAT and post-training quantization (PTQ). For example, QAT may be a method of acquiring quantization parameters by performing quantization and training together. Since QAT performs quantization and training together, the time taken to perform quantization may be relatively long. For example, PTQ may be a method of acquiring quantization parameters by performing quantization on an artificial intelligence model that has completed training. Since PTQ performs only quantization, the time taken to perform quantization according to PTQ may be relatively short. Therefore, an artificial intelligence model with a relatively large model size (e.g., a large language model (LLM)) may substantially only perform PTQ. To prevent latency in acquiring quantization parameters, an artificial intelligence model with a relatively large model size may not be used as a reference artificial intelligence model. For example, an AI model capable only of PTQ may not be used as a reference AI model for determining quantization parameters. For example, an AI model capable of QAT may be used as a reference AI model for determining quantization parameters. In one example, an AI model capable of QAT may be identified based on the model size. The server device (210) may identify the AI model as a model capable only of PTQ based on the identification that the size of the AI model exceeds a threshold size. The server device (210) may identify the AI model as a model capable of QAT based on the identification that the size of the AI model is less than a threshold size.The threshold size can be predefined to identify whether the artificial intelligence model is capable of only PTQ.
[0099] In one embodiment, the server device (210) can identify at least one artificial intelligence model by excluding an artificial intelligence model capable of updating and an artificial intelligence model capable only of PTQ from a plurality of artificial intelligence models according to the method described above. For example, at least one artificial intelligence model may be included in a list for identifying a reference artificial intelligence model.
[0100] In operation 702, a server device (210) according to one embodiment can determine quantization parameters based on a reference artificial intelligence model among at least one artificial intelligence model.
[0101] In one embodiment, the server device (210) can identify an artificial intelligence model having a minimum model size among at least one artificial intelligence model. The artificial intelligence model having a minimum model size may be a reference artificial intelligence model for determining quantization parameters among a plurality of artificial intelligence models. The model size may correspond to the number of operations of the corresponding artificial intelligence model and / or the size of the file of the corresponding artificial intelligence model. In one example, the artificial intelligence model having a minimum model size may be the artificial intelligence model having the smallest number of operations. In one example, the artificial intelligence model having a minimum model size may be the artificial intelligence model having a minimum file size.
[0102] In one embodiment, the server device (210) can perform quantization on a reference artificial intelligence model. The quantization method may include QAT and PTQ. QAT or PTQ may be used as the quantization method. The server device (210) may use either QAT or PTQ as the quantization method, taking into account physical time and / or the development environment (e.g., GPU (graphic processing unit), RAM (random access memory)). For an artificial intelligence model where QAT cannot be used due to consideration of physical time and / or the development environment, PTQ may be used as the quantization method. For example, the server device (210) may identify the quantization method based on the size of the reference artificial intelligence model. In one example, the server device (210) may perform quantization on the reference artificial intelligence model based on PTQ upon identifying that the model size exceeds a threshold size. By performing quantization on the reference artificial intelligence model based on PTQ, the time required to perform quantization on a reference artificial intelligence model having a large model size can be reduced. In one example, the server device (210) can perform quantization of a reference AI model based on QAT upon identifying that the model size is less than or equal to a threshold size. By performing quantization of the reference AI model based on QAT, more accurate quantization parameters can be obtained (or determined) for a reference AI model having a small model size. The threshold size can be predefined to determine the quantization method. After performing quantization, the server device (210) can generate a frozen graph for the reference AI model.
[0103] In operation 703, a server device (210) according to one embodiment can perform quantization of artificial intelligence models based on quantization parameters. For example, the server device (210) can perform quantization of the remaining artificial intelligence models, excluding the reference artificial intelligence model among a plurality of artificial intelligence models. After performing quantization of the remaining artificial intelligence models, the server device (210) can generate frozen graphs. In one example, a plurality of artificial intelligence models may include a first artificial intelligence model (e.g., the first artificial intelligence model (410) of FIG. 4b), a second artificial intelligence model (e.g., the second artificial intelligence model (420) of FIG. 4b), and a third artificial intelligence model (e.g., the third artificial intelligence model (430) of FIG. 4b). The output of the first artificial intelligence model and the output of the second artificial intelligence model may be used as inputs to the third artificial intelligence model. In one example, the first artificial intelligence model in the above-described ensemble model may be a reference artificial intelligence model. The server device (210) may obtain (or determine) output quantization parameters by performing quantization on the first artificial intelligence model. The server device (210) may perform quantization on the second artificial intelligence model such that the output quantization parameters of the second artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT. The server device (210) may perform quantization on the third artificial intelligence model such that the input quantization parameters of the third artificial intelligence model have the output quantization parameters of the first artificial intelligence model. Quantization can be performed on the model. The quantization method may be PTQ or QAT. However, this is merely an example to explain the structure of the ensemble model, and the present disclosure is not limited thereto. The structure of the ensemble model according to the present disclosure may have various forms in which the output of at least one artificial intelligence model among a plurality of artificial intelligence models is used as the input of another artificial intelligence model.For example, the output of the first artificial intelligence model may be used as the input of the second artificial intelligence model and the input of the third artificial intelligence model. In one example, the first artificial intelligence model in the aforementioned ensemble model may be a reference artificial intelligence model. The server device (210) may obtain (or determine) output quantization parameters by performing quantization on the first artificial intelligence model. The server device (210) may perform quantization on the second artificial intelligence model such that the input quantization parameters of the second artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT. The server device (210) may perform quantization on the third artificial intelligence model such that the input quantization parameters of the third artificial intelligence model have the output quantization parameters of the first artificial intelligence model. The quantization method may be PTQ or QAT.
[0104] In one embodiment, the server device (210) may generate frozen graphs for artificial intelligence models and then determine whether the accuracy of an ensemble model including the artificial intelligence models exceeds a threshold value. For example, the server device (210) may generate frozen graphs for a plurality of artificial intelligence models and then determine the accuracy of the ensemble model. For example, the accuracy of the ensemble model may correspond to a quantized signal-to-noise ratio (QSNR) or a custom metric. For example, a threshold value for evaluating accuracy may be predefined. For example, the server device (210) may perform operation 704 upon determining that the accuracy of the ensemble model exceeds a threshold value. In another example, the server device (210) may determine the next artificial intelligence model as the reference artificial intelligence model upon determining that the accuracy of the ensemble model is below a threshold value. For example, the next artificial intelligence model may be an artificial intelligence model having the second lowest model size among at least one artificial intelligence model identified in operation 701. The server device (210) can perform operations 702 and 703 on the next artificial intelligence model. The server device (210) can determine the accuracy of the ensemble model based on quantization parameters determined based on the next artificial intelligence model. The server device (210) can perform operation 704 upon determining that the accuracy of the ensemble model exceeds a threshold value. As described above, if the accuracy of the ensemble model is not satisfied, the server device (210) can select the next model in the list as the reference artificial intelligence model and then retry from sub-model quantization (e.g., operation 702).
[0105] In operation 704, a server device (210) according to one embodiment may transmit a file for a plurality of artificial intelligence models to an electronic device (101). For example, the server device (210) may generate a file for an ensemble model including a plurality of artificial intelligence models based on a determination that the accuracy of the ensemble model exceeds a threshold value. The server device (210) may transmit the generated file to the electronic device (101).
[0106] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.
[0107] A server device as described above may include a communication circuit. The server device may include a memory that stores instructions and includes one or more storage media. The server device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the server device may cause at least one artificial intelligence model to identify by excluding an updatable artificial intelligence model from a plurality of artificial intelligence models of an ensemble model. When the instructions are executed individually or collectively by the at least one processor, the server device may cause quantization parameters to be determined based on performing quantization on a reference artificial intelligence model among the at least one artificial intelligence model. When the instructions are executed individually or collectively by the at least one processor, the server device may cause quantization on one or more artificial intelligence models excluding the reference artificial intelligence model among the plurality of artificial intelligence models based on the quantization parameters. When the above instructions are executed individually or collectively by the at least one processor, the server device may cause the server device to transmit a file for the ensemble model including the plurality of artificial intelligence models to an electronic device.
[0108] For example, the plurality of artificial intelligence models may include a main artificial intelligence model and sub artificial intelligence models. The outputs of the sub artificial intelligence models may be used as inputs to the main artificial intelligence model.
[0109] For example, when the above instructions are executed individually or collectively by the at least one processor, the server device may cause the main artificial intelligence model to perform quantization of the main artificial intelligence model such that the input quantization parameters of the main artificial intelligence model have the quantization parameters. When the above instructions are executed individually or collectively by the at least one processor, the server device may cause the sub artificial intelligence models to perform quantization of the sub artificial intelligence models such that the output quantization parameters of the sub artificial intelligence models have the quantization parameters.
[0110] For example, when the above instructions are executed individually or collectively by the at least one processor, the server device may be caused to identify the reference artificial intelligence model having the minimum model size among the at least one artificial intelligence model.
[0111] For example, when the above instructions are executed individually or collectively by the at least one processor, the server device may be caused to identify the at least one artificial intelligence model by excluding the updateable artificial intelligence model from the plurality of artificial intelligence models of the ensemble model and identifying the available artificial intelligence model capable of quantization-aware training (QAT).
[0112] For example, when the above instructions are executed individually or collectively by the at least one processor, the server device may cause the reference artificial intelligence model to perform quantization based on a quantization method identified based on the size of the reference artificial intelligence model. The quantization method may be QAT or PTQ (post training quantization).
[0113] For example, when the instructions are executed individually or collectively by the at least one processor, the server device may cause the server device to determine whether an evaluation metric for the ensemble model including the plurality of artificial intelligence models exceeds a threshold value. When the instructions are executed individually or collectively by the at least one processor, the server device may cause the server device to transmit the file for the ensemble model to the electronic device upon the determination that the evaluation metric for the ensemble model exceeds the threshold value. When the instructions are executed individually or collectively by the at least one processor, the server device may cause the server device to identify a second reference artificial intelligence model among the at least one artificial intelligence model upon the determination that the evaluation metric for the ensemble model is below the threshold value. When the above instructions are executed individually or collectively by the at least one processor, the server device may cause the server device to perform quantization of the plurality of artificial intelligence models based on quantization parameters determined based on performing quantization of the second reference artificial intelligence model.
[0114] For example, when an update is performed on the above-mentioned updateable artificial intelligence model, an update may not be performed on at least one of the above-mentioned artificial intelligence models.
[0115] For example, the above quantization parameters may include scale and offset.
[0116] A method performed by a server device as described above may include an operation of identifying at least one artificial intelligence model by excluding an updatable artificial intelligence model from a plurality of artificial intelligence models of an ensemble model. The method may include an operation of determining quantization parameters based on performing quantization on a reference artificial intelligence model among the at least one artificial intelligence model. The method may include an operation of performing quantization on one or more artificial intelligence models excluding the reference artificial intelligence model among the plurality of artificial intelligence models based on the quantization parameters. The method may include an operation of transmitting a file for the ensemble model including the plurality of artificial intelligence models to an electronic device.
[0117] For example, the plurality of artificial intelligence models may include a main artificial intelligence model and sub artificial intelligence models. The outputs of the sub artificial intelligence models may be used as inputs to the main artificial intelligence model.
[0118] For example, the above method may include an operation of performing quantization on the main artificial intelligence model such that the input quantization parameters of the main artificial intelligence model have the quantization parameters. The above method may include an operation of performing quantization on the sub artificial intelligence models such that the output quantization parameters of the sub artificial intelligence models have the quantization parameters.
[0119] For example, the above method may include the operation of identifying the reference artificial intelligence model having a minimum model size among the at least one artificial intelligence model.
[0120] For example, the above method may include the operation of identifying at least one artificial intelligence model by excluding the updateable artificial intelligence model from the plurality of artificial intelligence models of the ensemble model and identifying the quantization-aware (QAT) training-available artificial intelligence model.
[0121] For example, the above method may include an operation of performing quantization on the reference artificial intelligence model based on a quantization method identified based on the size of the reference artificial intelligence model. The quantization method may be QAT or PTQ.
[0122] For example, the above method may include an operation of determining whether an evaluation metric for an ensemble model including the plurality of artificial intelligence models exceeds a threshold value. The above method may include an operation of transmitting the file for the ensemble model to the electronic device in accordance with the determination that the evaluation metric for the ensemble model exceeds the threshold value. The above method may include an operation of identifying a second reference artificial intelligence model among the at least one artificial intelligence model in accordance with the determination that the evaluation metric for the ensemble model is less than or equal to the threshold value. The above method may include an operation of performing quantization on the plurality of artificial intelligence models based on quantization parameters determined based on performing quantization on the second reference artificial intelligence model.
[0123] For example, when an update is performed on the above-mentioned updateable artificial intelligence model, an update may not be performed on at least one of the above-mentioned artificial intelligence models.
[0124] For example, the above quantization parameters may include scale and offset.
[0125] The electronic device described above may include a communication circuit. The electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include at least one processor that includes a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause a first output according to a first input by using a first artificial intelligence model among a plurality of artificial intelligence models stored in the memory. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause a second output according to a second input by using a second artificial intelligence model among the plurality of artificial intelligence models. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the output quantization parameters of the first artificial intelligence model and the output quantization parameters of the second artificial intelligence model to refrain from reading from the memory in order to perform de-quantization for the first output and the second output. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the input quantization parameters for the third artificial intelligence model among the plurality of artificial intelligence models to refrain from reading from the memory. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the third artificial intelligence model among the plurality of artificial intelligence models to generate an output based on the first output and the second output.
[0126] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to receive an update file for a first artificial intelligence model among an ensemble model comprising the plurality of artificial intelligence models from a server device. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the first artificial intelligence model to be updated based on the update file. Updates for the second artificial intelligence model and the third artificial intelligence model may not be performed.
[0127] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.
[0128] For one or more embodiments, at least one of the components described in one or more of the prior art drawings may be configured to perform one or more operations, techniques, processes and / or methods as described in the present disclosure. For example, a processor (e.g., a baseband processor) described in the present disclosure in relation to one or more of the prior art drawings may be configured to operate according to one or more examples described in the present disclosure. As another example, circuits associated with user equipment (UE), a base station, a network element, etc., as described above in relation to one or more of the prior art drawings may be configured to operate according to one or more examples described herein.
[0129] Any of the embodiments described above may be combined with any other embodiment (or combination of embodiments) unless otherwise explicitly stated. The foregoing description of one or more embodiments is for illustrative and explanatory purposes only, and is not intended to limit or exhaust the scope of the embodiments in the exact form disclosed. Modifications and variations are possible in light of the foregoing teachings or may be obtained from the practice of various embodiments.
[0130] The electronic devices according to the various embodiments disclosed in this document may be of various forms. The electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, electronic devices, or consumer electronics. The electronic devices according to the embodiments of this document are not limited to the devices described above.
[0131] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0132] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0133] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0134] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0135] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In a server device, Communication circuit; Memory for storing instructions and including one or more storage media; and It includes at least one processor comprising a processing circuit, and When the above instructions are executed individually or collectively by the at least one processor, the server device, Identifying at least one artificial intelligence model by excluding an updatable artificial intelligence model from multiple artificial intelligence models of an ensemble model, and Based on performing quantization on a reference AI model among the above at least one AI model, quantization parameters are determined, and Based on the above quantization parameters, quantization is performed on one or more artificial intelligence models among the plurality of artificial intelligence models, excluding the reference artificial intelligence model, and Causing to transmit a file for the ensemble model including the plurality of artificial intelligence models to an electronic device Server device.
2. In Paragraph 1, The above plurality of artificial intelligence models include a main artificial intelligence model and sub artificial intelligence models, and The outputs of the above sub-AI models are used as inputs to the above main AI model, Server device.
3. In Paragraph 2, When the above instructions are executed individually or collectively by the at least one processor, the server device, Quantization is performed on the main artificial intelligence model such that the input quantization parameters of the main artificial intelligence model have the quantization parameters, and Causing the output quantization parameters of the above sub-artificial intelligence models to have the above quantization parameters, to perform quantization for the above sub-artificial intelligence models, Server device.
4. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the server device, Causing to identify the reference artificial intelligence model having the minimum model size among the above at least one artificial intelligence model, Server device.
5. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the server device, Causing to identify at least one artificial intelligence model by excluding the updateable artificial intelligence model from the plurality of artificial intelligence models of the above ensemble model and identifying the QAT (quantization aware training) available artificial intelligence model, Server device.
6. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the server device, Causing to perform quantization of the reference artificial intelligence model based on a quantization method identified based on the size of the reference artificial intelligence model, and The above quantization method is QAT or PTQ (post-training quantization), Server device.
7. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the server device, Determining whether an evaluation metric for the ensemble model including the plurality of artificial intelligence models exceeds a threshold value, Based on the determination that the evaluation metric for the ensemble model exceeds the threshold value, the file for the ensemble model is transmitted to the electronic device, and Based on the determination that the evaluation metric for the above ensemble model is below the threshold value, a second reference artificial intelligence model is identified among the at least one artificial intelligence model, and Causing to perform quantization of the plurality of artificial intelligence models based on quantization parameters determined based on performing quantization of the second reference artificial intelligence model, Server device.
8. In Paragraph 1, When an update is performed on the above-mentioned updateable artificial intelligence model, updates are not performed on the artificial intelligence models among the above-mentioned plurality of artificial intelligence models, excluding the above-mentioned updateable artificial intelligence model. Server device.
9. In Paragraph 1, The above quantization parameters include scale and offset, Server device.
10. In a method performed by a server device, An operation of identifying at least one artificial intelligence model by excluding an updatable artificial intelligence model from multiple artificial intelligence models of an ensemble model; An operation to determine quantization parameters based on performing quantization on a reference artificial intelligence model among the above at least one artificial intelligence model; Based on the above quantization parameters, an operation of performing quantization on one or more artificial intelligence models excluding the reference artificial intelligence model among the plurality of artificial intelligence models; and The operation of transmitting a file for the ensemble model including the plurality of artificial intelligence models to an electronic device, method.
11. In Paragraph 10, The above plurality of artificial intelligence models include a main artificial intelligence model and sub artificial intelligence models, and The outputs of the above sub-AI models are used as inputs to the above main AI model, method.
12. In Paragraph 11, An operation to perform quantization on the main artificial intelligence model such that the input quantization parameters of the main artificial intelligence model have the quantization parameters; and further comprising an operation of performing quantization for the sub-artificial intelligence models such that the output quantization parameters of the sub-artificial intelligence models have the quantization parameters. method.
13. In Paragraph 10, Further including the operation of identifying the reference artificial intelligence model having a minimum model size among the at least one artificial intelligence model, method.
14. In paragraph 10, the operation of identifying at least one artificial intelligence model is, The method comprises the operation of identifying at least one artificial intelligence model by excluding the updateable artificial intelligence model from the plurality of artificial intelligence models of the ensemble model and identifying the quantization-aware (QAT) training-available artificial intelligence model. method, 15. In Paragraph 10, the operation of performing quantization on the above-mentioned reference artificial intelligence model is, The method includes an operation to perform quantization on the reference artificial intelligence model based on a quantization method identified based on the size of the reference artificial intelligence model, and The above quantization method is QAT or PTQ, method.
Citation Information
Patent Citations
Model ensemble generation
JP7119751B2
Photomask and manufacturing method thereof
KR1020240013047A
Weight-based ensemble model selection system and method
KR102535010B1
Method for evaluating an artificial neural network model performance and system using the same
US12106209B1
KR20210055992A