Method and device for training and inference of artificial intelligence related to data quantization in wireless communication system
The quantization error-optimized AI model addresses the challenge of high power and complexity in wireless communication systems by optimizing quantization noise, enhancing performance and efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
Existing AI models for wireless communication systems face challenges in achieving low power and low complexity while maintaining performance due to quantization errors, leading to increased computational complexity and power consumption.
A quantization error-optimized AI model is developed with bit interpretation and a training method that involves determining maximum and minimum values of output nodes, updating quantization units, and applying quantization noise to optimize the model for low power and low complexity.
The proposed method improves device performance by optimizing quantization noise in AI learning and inference, reducing power consumption and complexity without the need for separate optimization processes.
Smart Images

Figure KR2025017740_07052026_PF_FP_ABST
Abstract
Description
Artificial intelligence learning and inference methods and devices related to data quantization in wireless communication systems
[0001] The present disclosure relates to an artificial intelligence learning and inference method and apparatus related to quantization in a wireless communication system.
[0002] With the development of digital technology, electronic devices are being provided in various forms, such as smartphones, tablet PCs, or PDAs.
[0003] As artificial intelligence technology advances, electronic devices can provide various artificial intelligence services by applying artificial intelligence technology. Based on artificial intelligence technology and voice recognition technology, electronic devices can provide artificial intelligence services that process tasks requested by the user and offer software configurations (e.g., services, functions, or programs) that provide services specialized for the user (e.g., customized information based on user voice commands).
[0004] Machine learning is a field related to artificial intelligence that develops algorithms and technologies enabling computers to learn. Deep learning refers to a set of machine learning algorithms that attempt a high level of abstraction (the task of extracting only the essential content from a large amount of complex data) through a combination of non-linear transformation techniques.
[0005] Meanwhile, as interest in AI-applied on-device products has recently increased, there is a demand for technological advancements in the research and design of low-power, low-complexity AI inference engines.
[0006] The information described above may be provided as related art for the purpose of aiding understanding of this document. None of the foregoing is to be claimed as prior art related to this document, nor is it to be used to determine prior art.
[0007] The present disclosure proposes a quantization error-optimized AI model capable of bit interpretation to have low power and low complexity, and proposes a training method for said AI model, an inference method for applying said AI model, and an apparatus.
[0008] A method for learning an artificial intelligence model related to the quantization of an electronic device in a wireless communication system according to one embodiment of the present disclosure comprises: determining a maximum value and a minimum value of the output value of a first node included in the artificial intelligence model related to the quantization; updating the maximum value and the minimum value; calculating a quantization unit based on the updated maximum value and the minimum value; and updating quantization noise applied to the artificial intelligence model based on the quantization unit.
[0009] According to one embodiment of the present disclosure, a storage medium storing at least one computer-readable instruction, wherein the at least one instruction causes the electronic device to perform at least one operation when executed by at least a part of at least one processor (120) of the electronic device, and the at least one operation includes: an operation of determining a maximum value and a maximum value of an output value of a first node included in a quantization-related artificial intelligence model; an operation of updating the maximum value and a minimum value; an operation of calculating a quantization unit based on the updated maximum value and a minimum value; and an operation of updating quantization noise applied to the artificial intelligence model based on the quantization unit.
[0010] According to one embodiment of the present disclosure, an electronic device comprises at least one processor (120); and a memory (130) for storing at least one instruction, wherein the at least one instruction causes the electronic device to perform at least one operation when executed by at least a part of the at least one processor (120), and the at least one operation includes: an operation of determining a maximum value and a maximum value of an output value of a first node included in a quantization-related artificial intelligence model; an operation of updating the maximum value and a minimum value; an operation of calculating a quantization unit based on the updated maximum value and a minimum value; and an operation of updating quantization noise applied to the artificial intelligence model based on the quantization unit.
[0011] Through the present disclosure, the performance of the device can be improved by performing quantization noise-optimized AI learning and inference.
[0012] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0013] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment of the present disclosure.
[0014] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment of the present disclosure.
[0015] FIG. 3 is a diagram illustrating a floating-point artificial intelligence (AI) model in a network node according to one embodiment of the present disclosure.
[0016] FIG. 4 is a diagram showing an example of a dynamic region of a signal and a B-bit quantization model according to one embodiment of the present disclosure.
[0017] FIG. 5 illustrates an example of an equivalent quantization noise model according to one embodiment of the present disclosure.
[0018] FIG. 6 is a flowchart illustrating a quantization recognition learning operation according to one embodiment of the present disclosure.
[0019] FIG. 7 is a flowchart illustrating the optimization operation of an artificial intelligence inference model considering quantization noise according to one embodiment of the present disclosure.
[0020] FIG. 8 is a flowchart illustrating a quantization optimal learning operation according to one embodiment of the present disclosure.
[0021] FIG. 9 is a drawing showing an example of a fixed-point AI model including N nodes according to one embodiment of the present disclosure.
[0022] FIG. 10 is a drawing showing an example of a simple equivalent quantization noise AI model according to one embodiment of the present disclosure.
[0023] FIG. 11 is a diagram illustrating an example of an adaptive equivalent quantization noise AI model according to one embodiment of the present disclosure.
[0024] FIG. 12 is a block diagram of a quantization optimal AI model according to one embodiment of the present disclosure.
[0025] FIG. 13 is a detailed block diagram of the nth node included in a quantization optimal AI model according to one embodiment of the present disclosure.
[0026] FIG. 14 is a diagram illustrating a method for updating the maximum value (Max) and minimum value (Min) of the nth node included in a quantization optimal AI model according to one embodiment of the present disclosure.
[0027] FIG. 15 is a diagram illustrating a method of applying quantization noise of the nth node included in a quantization optimal AI model according to one embodiment of the present disclosure.
[0028] FIG. 16 is a flowchart illustrating the learning operation of an artificial intelligence model related to the quantization of an electronic device according to one embodiment of the present disclosure.
[0029] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0030] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0031] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0032] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0033] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0034] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0035] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0036] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0037] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0038] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0039] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0040] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0041] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0042] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0043] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0044] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0045] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0046] A communication module (transmitter / receiver) (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of a processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0047] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0048] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0049] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0050] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0051] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0052] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0053] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0054] As used in various embodiments of this document, the term “module” may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0055] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0056] In the following description, the electronic device may include a receiving device that receives a signal (or data) from a transmitting device of a wireless communication system. For example, the electronic device may include a base station that receives a signal (or data) from a terminal. As an example, the base station may be at least one of a Node B, a BS (base station), an eNB (eNode B), or a gNB (gNode B). For example, the electronic device may include a terminal that receives a signal (or data) from a base station. As an example, the terminal may be at least one of a UE (user equipment), an MS (mobile station), a cellular phone, a smartphone, a computer, or a multimedia system capable of performing communication functions.
[0057] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment of the present disclosure.
[0058] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment of the present disclosure.
[0059] According to one embodiment, the electronic device (101) may include, and / or execute a frontend module (210) and / or a service processing module (220). The frontend module (210) and / or the service processing module (220) may be executed, for example, by a processor (120) or may be included as at least part of the processor (120) or other entity. At least some of the operations performed by the frontend module (210) and / or the service processing module (220) in the present disclosure may be understood as being performed, for example, by the processor (120) and / or other entity under the control of the processor (120).
[0060] The frontend module (210) can perform at least one operation for exchanging data with, for example, external electronic devices (106a, 106b, 106c, ..., 106n). For example, the frontend module (210) can provide data that can configure a user interface (UI) for inputting user input from the external electronic devices (106a, 106b, 106c, ..., 106n). The frontend module (210) can provide processing for user requests to the service processing module (220). The service processing module (220) can perform a service using the user request and may be named a backend module. The service processing module (220) can provide a response corresponding to the user request to the frontend module (210). The frontend module (210) can provide the response received from the service processing module (220) to an external electronic device (106a, 106b, 106c, ..., 106n).
[0061] The service processing module (220) may include, for example, a user request verification module (221), an optimization module (222), an AI model management module (224), a policy management module (226) and / or a service execution module (227).
[0062] The user request verification module (221) can verify information associated with user requests provided by external electronic devices (106a, 106b, 106c, ..., 106n). Information associated with user requests may be expressed, for example, as the number of user requests and / or the size of user requests over a certain period, but there are no limitations. For example, the user request verification module (221) may count the number of user requests provided by external electronic devices (106a, 106b, 106c, ..., 106n) and / or monitor the size of content included in the user requests (e.g., text and / or graphic objects, but there are no limitations), but there are no limitations on the type of information associated with user requests and / or the verification method.
[0063] The optimization module (222) can provide at least one optimal number of instances for each of at least one AI model associated with the service.
[0064] The AI model management module (224) can store and / or manage AI models associated with the service (e.g., may include, but is not limited to, adding, deleting, and / or updating).
[0065] The policy management module (226) may, for example, store and / or manage allowable response times (e.g., may include, but is not limited to, adding, deleting, and / or updating). For example, the management device (104) may verify (e.g., receive or determine) at least one input for determining allowable response times and provide it to the electronic device (101). For example, an administrator may input information regarding allowable response times for the corresponding service into the management device (104), but this is exemplary and there is no limitation on the way allowable response times are verified.
[0066] The service execution module (227) can execute each of at least one instance group (231, 232, 233) corresponding to at least one AI model associated with the service, for example. The service execution module (227) can execute at least one instance group (231, 232, 233) corresponding to at least one AI model according to the optimal number of instances provided by the optimization module (222), for example. The service execution module (227) can process user requests based on the executed at least one instance group (231, 232, 233) and can provide a response according to the processing result to an external electronic device (106a, 106b, 106c, ..., 106n) through the frontend module (210). According to one embodiment, if there are multiple instances being executed, user requests can be distributed and processed by the multiple instances.
[0067] The error due to quantization can be modeled as follows using the 1-bit 6dB rule. Assuming an AI node with floating-point numbers as follows, if the output value at node A obtained through learning is x, the range of maximum and minimum values and statistical characteristics for x for each node can be obtained using floating-point model simulation.
[0068] FIG. 3 is a diagram illustrating a floating-point artificial intelligence (AI) model in a network node according to one embodiment of the present disclosure.
[0069] FIG. 3 illustrates a structural diagram of a floating-point model. Referring to FIG. 3, the output passing through the floating-point AI model (300) can be represented as data X. The data X can be quantized into a bit size B as shown in FIG. 4. In this case, the entire data 2B -1 It can be quantized into steps, and in this case, each quantization unit It can be expressed as shown in mathematical formula 1 below.
[0070] [Mathematical Formula 1]
[0071]
[0072] Here, am.
[0073] Quantization noise e is uniformly distributed at the quantization unit, and when applying quantization, a round method (e.g., a method of quantizing by rounding data values) (i.e., maximum quantization noise Assuming the application of (generated by), quantization noise power It can be derived as shown in mathematical formula 2 below.
[0074] [Mathematical Formula 2]
[0075]
[0076] Meanwhile, signal power signal-to-quantization noise ratio (SQNR) Defined as Assuming this, if expressed in dB form, 1 bit as shown in Mathematical Formula 3 below It can be expressed as 6dB.
[0077] [Mathematical Formula 3]
[0078]
[0079] In the above mathematical formula 3, when applying quantization, the floor method (for example, a method of quantizing by rounding data values) (i.e., quantization noise Assuming that a uniform distribution is applied between them, it can be expressed as Equation 4 below.
[0080] [Mathematical Formula 4]
[0081]
[0082] The above mathematical formula 4 can be rearranged into a quantization noise formula as mathematical formula 5 depending on the quantization method.
[0083] [Mathematical Formula 5]
[0084]
[0085] Here, if the quantization method is round, C can be defined as 10.8, and if it is floor, C can be defined as 4.77.
[0086] The above mathematical equation 5 indicates that the SQNR due to quantization noise increases by approximately 6 dB times the set bit size B applied to quantization, and the dynamic range of the signal applied to quantization It may imply that it is inversely proportional to. Additionally, the above mathematical equation 5 indicates that if the dynamic range (DR) of the signal for each AI node can be measured, the quantization noise can be mathematically modeled according to the size of the number of bits B that sets the node. This implies that when applying a B-bit quantization model like that of Fig. 4 to the floating-point model of Fig. 3, it can be mathematically modeled as shown in Fig. 5.
[0087] FIG. 5 illustrates an example of an equivalent quantization noise model according to one embodiment of the present disclosure.
[0088] Referring to FIG. 5, quantization noise (510) is applied to data that has passed through a specific node A of a floating AI model to derive an output (520) that reflects an equivalent quantization noise error. In this case, the quantization noise (510) can be applied as shown in Equation 6 below.
[0089] [Mathematical Formula 6]
[0090]
[0091] In quantization-aware training (QAT), a training method that considers quantization noise can be applied as follows.
[0092] (1) Development of a floating-point AI model without quantization
[0093] (2) Apply a quantization model with a specific scale and bit saturation (fake quantization) to the hyperparameters (e.g., weights) of the forward pass and the activation layer, and train the AI model without applying the quantization model to the back-pass.
[0094] (3) Use this to select bits suitable for the inference model
[0095] (4) Fixed-point inference engine applied
[0096] Hereinafter, FIGS. 6 and 7 specifically explain an example of the above-mentioned quantization recognition learning method.
[0097] FIG. 6 is a flowchart illustrating a quantization recognition learning operation according to one embodiment of the present disclosure.
[0098] Referring to FIG. 6, for quantization recognition learning, forward pass (640) and backward propagation (650) operations can be performed. Forward pass (640) may refer to calculating and storing variables sequentially from the input layer to the output layer of the neural network of the AI model. Backward propagation (650) refers to a method of calculating the gradients for the parameters of the neural network of the AI model. In general, back propagation may refer to calculating and storing the gradients of the intermediate variables and parameters of the objective function associated with each layer of the neural network in the order from the output layer to the input layer.
[0099] More specifically, in the case of a forward pass (640), a quantization effect (600) with a fixed scale is applied to the weights, which are parameters of the AI model, and a matrix operation (610) can be performed on the input values. Afterwards, a bias (615), which is a parameter of the AI model, is applied, and an activation layer (620) and a quantization effect (630) with a fixed scale are applied to calculate output values up to the output layer. In the case of backward propagation (650), the AI model can be trained without applying the quantization effect (600, 630) with a fixed scale that was applied in the forward pass (640).
[0100] FIG. 7 is a flowchart illustrating the optimization operation of an artificial intelligence inference model considering quantization noise according to one embodiment of the present disclosure.
[0101] Referring to FIG. 7, the AI model can be optimized for the output values of the floating-point model nodes (700, 710, ..., 720) and the fixed-point model nodes (730, 740, ... 750) through the score comparison unit (760) and the best-K model selection unit (770).
[0102] That is, regarding the selection of bits suitable for the AI model training and inference model mentioned above, a method for optimizing quantization noise can be applied by using the best K selection method, which uses the performance difference between the result of the floating-point model and the quantized fixed-point model as an indicator. The Best K method is a method of finding K fixed-point models with a small performance difference from the floating-point model by listing the results of quantization performed for each node in order as shown in Figure 7. However, as the number of nodes increases, the utility of this method may decrease because the number of quantization selections per node increases, and the cost of optimizing through the construction of fixed-point models and performance comparison with floating-point models becomes high.
[0103] For example, if the number of AI nodes is N and there are three bit selections available for each AI node—4-bit, 8-bit, and 16-bit—then 3N selections occur, and a process of comparing performance with floating-point models is required, which can increase development time and costs. Furthermore, since the hyperparameters selected in the above manner are values optimized during the floating-point model process, they may not be considered optimal hyperparameters when the model is changed to a fixed-point model.
[0104] Furthermore, since the scale value of the quantization process is used as a fixed value per node, data cannot be found instantaneously. Because a floating model must be constructed in advance and appropriate quantization scale values applied to each node after analyzing node-specific statistics, an optimization process for the quantization scale is required when the training set changes.
[0105] FIG. 8 is a flowchart illustrating a quantization optimal learning operation according to one embodiment of the present disclosure.
[0106] Referring to FIG. 8, for quantization recognition learning, forward pass (850) and backward propagation (860) operations can be performed.
[0107] More specifically, in the case of a forward pass (850), a quantization effect (800) can be applied to the weights, which are parameters of the AI model, using a data-adaptive scale, and a matrix operation (810) can be performed on the input values. Subsequently, a quantization effect (820) can be applied using a fixed scale, a bias (825) can be applied to the biases, which are parameters of the AI model, and an activation layer (830) can be used to calculate output values up to the output layer using a fixed scale quantization effect (840). In this case, the data-adaptive scale can be adaptively changed and applied whenever the quantization effects (800, 820, and 840) are applied.
[0108] In the case of backward propagation (860), the AI model can be trained without applying the quantization effects (800, 820 and 840) that were applied in the forward pass (850).
[0109] The present disclosure proposes a bit-analyzable quantization noise optimization AI modeling and training method to address the increased cost in the process of optimizing the quantization inference engine in an AI model and the hyperparameter mismatch occurring during the quantization process, and provides a mathematical model for training this method.
[0110] By utilizing an AI model that considers the mathematical model of quantization noise provided in this disclosure, the quantization noise during the training process can be modeled as data-derived equivalent quantization noise or simple equivalent quantization noise. This allows the quantization model to be applied adaptively to the training set in the forward pass, thereby enabling the AI model to be optimally trained for the quantization noise. By applying this, hyperparameters and per-node bits suitable for a fixed-point inference engine can be derived.
[0111] In addition, without a separate quantization optimization process, an inference model or engine can be created by applying the previously optimized hyperparameters and node-specific bits.
[0112] FIG. 9 is a drawing showing an example of a fixed-point AI model including N nodes according to one embodiment of the present disclosure.
[0113] Referring to FIG. 9, the fixed-point model may include N nodes (910, 920, ..., 950). As illustrated in FIG. 9, for each node, the number of bits is B1, ..., BN, and the quantization unit is When considering quantization, an equivalent quantization noise model can be considered as shown in Fig. 10.
[0114] FIG. 10 is a drawing showing an example of a simple equivalent quantization noise AI model according to one embodiment of the present disclosure.
[0115] The modeling illustrated in FIG. 10 may include a method of modeling quantization noise occurring in a fixed-point model into a floating-point model (1010, 1020, ..., 1050) and additive Gaussian noise (N1, N2, ..., NN). FIG. 10 illustrates a method of simply modeling with additive Gaussian in a form derived from FIG. 5 illustrated above.
[0116] Meanwhile, in the equivalent model illustrated in Fig. 10, Gaussian noise is modeled independently regardless of the input and output data for each node; therefore, to obtain optimized hyperparameters, specific training data may need to be obtained through multiple independent trial simulations so that it can sufficiently converge to the noise. For example, when N trials are required for the noise and the number of training data sets is M, M×N trials of training may be required. In this case, the larger the number of N trials, the more statistically accurate the modeling of the noise can become.
[0117] Unlike the method of modeling with simple equivalent quantization noise as shown in FIG. 10, FIG. 11 below describes a data-derived equivalent quantization noise model that constructs an equivalent model by applying quantization modeling suitable for the data format to each node of the AI model.
[0118] FIG. 11 is a diagram illustrating an example of an adaptive equivalent quantization noise AI model according to one embodiment of the present disclosure.
[0119] Referring to Fig. 11, the adaptive equivalent quantization noise AI model consists of floating model nodes (1110, 1120, ..., 1150) and - Quantization nodes (1115, 1125, ..., 1155) can be configured by intersecting.
[0120] That is, in Fig. 11, the quantization process is applied during the data scaling and data rounding (or flooring) processes, and in order to model this - Illustrate the implementation of the nth node by adding a quantization process. - An example of the quantization process expressed in a formula is mathematical equation 7 below.
[0121] [Mathematical Formula 7]
[0122]
[0123] Here and and can represent the maximum and minimum values of the n-th node data up to the j-th training set, respectively. Additionally, rof(·) can include a round or floor function to convert floating-point data to integers. The choice between the round and floor functions may depend on the fixed-point conversion function applied per node.
[0124] The above - During the quantization process, the maximum and minimum values are found according to the data format of the training set, and the quantization unit (scale) value is adaptively calculated to match the number of quantization bits applied to the node and reflected in the training.
[0125] FIG. 12 is a block diagram of a quantization optimal AI model according to one embodiment of the present disclosure.
[0126] Figure 12 illustrates a configuration diagram of a quantized optimal AI model including N nodes.
[0127] The configuration diagram for the nth node of the quantization optimal AI model illustrated in Fig. 12 is as shown in Fig. 13. That is, the operation at the nth node illustrated in Fig. 13 can be repeated N times to train the quantization optimal AI model.
[0128] Referring to Fig. 13, the output value of the j-th training set of the floating model of the n-th node Assuming that, in the Finding Max, Min block (1300) The maximum value of ( ) and minimum (min) value( ) can be found. An example of this expressed in a formula is mathematical formula 8 below.
[0129] [Mathematical Formula 8]
[0130]
[0131] Afterwards, in the Update Max, Min at node n block (1310), the maximum and minimum values can be set as shown in the following mathematical formula 9.
[0132] [Mathematical Formula 9]
[0133]
[0134] Here can represent the maximum value of the n-th node up to the j-1th training set. In addition, the above can represent the minimum value of the n-th node up to the j-1th training set. This mathematical process is diagrammed as shown in Figure 14.
[0135] Using the method illustrated in Fig. 14, the quantization unit (scale) of the n-th node using the updated maximum and minimum values at the n-th node It can be calculated using the above mathematical formula 1.
[0136] Quantization unit (scale) calculated using the above mathematical formula 1 By applying - In the DRFM (divide-round or floor-multiply) block (1330), the quantization process can be modeled as shown in the formula in Fig. 15.
[0137] By applying a quantization model as shown in Fig. 13 and configuring a total of N AI models as shown in Fig. 12, and performing training, the quantization noise is gradually updated while updating the maximum and minimum values for each node during each training process, and the training process can be performed to derive hyperparameters with the final quantization level applied by finally completing training on the entire training set.
[0138] By applying the quantization method proposed in this disclosure, quantization noise-optimized AI learning can be performed, and through this, hyperparameters of an AI model incorporating the proposed quantization model can be derived. Since the hyperparameters are applied as hyperparameters of a fixed-point inference engine without a separate bit optimization process, no additional optimization process is required during development, and the effect of low power consumption and low complexity can be achieved through the learning of the quantization noise-optimized AI model.
[0139] Below, an example of a method for applying the method proposed in the present disclosure to bit optimization is described.
[0140] Method 1) Optimization of optimal quantization bits per node through learning including selectable quantization bit candidates per node
[0141] The above method 1 comprises a set of candidate quantization bits for N nodes. Assuming, the set Since selection is possible per node from the K quantization bits, per training vector It is a method that minimizes the cost of bit usage and the cost function of the AI network by performing learning through the selection of.
[0142] Applying Method 1 allows it to be used for selecting the best quantization bits for each of N nodes. Here, the network's cost function Assuming..., the new cost function It can be defined as in mathematical formula 10.
[0143] [Mathematical Formula 10]
[0144]
[0145] Here, is a set It is a function representing the size of, and represents the quantization bit size selected at the nth node and , can mean the sum of the number of quantization bits available across all nodes. Generally It can be. As an example, In this case, it may mean proceeding with network training without quantization bit optimization.
[0146] Method 2) Quantization noise optimization and fine-tuning based on specific node bit interpretation
[0147] The quantization noise applied network proposed in this disclosure can be utilized in a method of optimizing quantization bits by determining the node with the greatest bit influence during the process of replacing each node with a floating model. Alternatively, the quantization model of this disclosure can be utilized in a fine-tuning process to find suitable bits for a specific node after training in the form of a floating model and designing bits for all nodes. The cost function applied in these processes can be modified as shown in Equation 11 below by applying the bit selection process of a specific single node to the cost function described in Method 1.
[0148] [Mathematical Formula 11]
[0149]
[0150] FIG. 16 is a flowchart illustrating the learning operation of an artificial intelligence model related to the quantization of an electronic device according to one embodiment of the present disclosure.
[0151] In step 1600, the electronic device can determine the maximum value and maximum value of the output value of the first node included in the quantization-related artificial intelligence model.
[0152] In step 1610, the electronic device can update the maximum and minimum values.
[0153] In step 1620, the electronic device can calculate the quantization unit based on the updated maximum and minimum values.
[0154] In step 1630, the electronic device can update the quantization noise applied to the artificial intelligence model based on the quantization unit.
[0155] In one embodiment, the electronic device can perform training on an artificial intelligence model to which the updated quantization noise is applied. In one embodiment, the electronic device can determine parameters for the artificial intelligence model. In one embodiment, the electronic device can perform inference using the artificial intelligence model based on the determined parameters.
[0156] In one embodiment, the output value of a floating-point model included in the first node may be included. In one embodiment, the operation of updating the quantization noise may include: dividing the output values of the floating-point model by the quantization unit; calculating a first value by rounding the output values of the floating-point model to the nearest value among the values corresponding to the quantization unit or by rounding down to the values corresponding to the quantization unit; and multiplying the first value by the quantization unit.
[0157] In one embodiment, the electronic device can calculate the output value of the second node of the artificial intelligence model for an input value to which the updated quantization noise has been applied.
[0158] In one embodiment, the operation of updating the maximum and minimum values includes the operation of updating the maximum value, and the operation of updating the maximum value is such that the maximum value of the output value of the first node is In this case, the first maximum value up to the j-th training data set of the first node ( ) is the second maximum value up to the j-1th training data set of the first node ( If it is greater than ), It may include an action to update to.
[0159] In one embodiment, the operation of updating the maximum and minimum values includes the operation of updating the minimum value, and the operation of updating the minimum value is such that the minimum value of the output value of the first node is In this case, the first minimum value up to the j-th training data set of the first node ( ) is the second minimum value up to the j-1th training data set of the first node ( If it is smaller than ), It may include an action to update to.
[0160] A storage medium for storing at least one computer-readable instruction according to one embodiment of the present disclosure, wherein the at least one instruction causes the electronic device to perform at least one operation when executed by at least a part of at least one processor (120) of the electronic device, and the at least one operation is: an operation of determining a maximum value and a maximum value of an output value of a first node included in a quantization-related artificial intelligence model;
[0161] The method includes: an operation to update the maximum and minimum values; an operation to calculate a quantization unit based on the updated maximum and minimum values; and an operation to update the quantization noise applied to the artificial intelligence model based on the quantization unit.
[0162] An electronic device according to one embodiment of the present disclosure comprises at least one processor (120); and a memory (130) for storing at least one instruction, wherein the at least one instruction causes the electronic device to perform at least one operation when executed by at least part of the at least one processor (120), and the at least one operation includes: an operation of determining a maximum value and a maximum value of an output value of a first node included in a quantization-related artificial intelligence model; an operation of updating the maximum value and a minimum value; an operation of calculating a quantization unit based on the updated maximum value and a minimum value; and an operation of updating quantization noise applied to the artificial intelligence model based on the quantization unit.
[0163] The present disclosure proposes a method and apparatus for a quantization-optimized AI learning and inference model capable of bit interpretation, an AI model learning model that adaptively reflects quantization errors, optimizes quantization bits and AI learning parameters for each node through learning, and an inference model that applies the derived node-specific bit information and learning parameters.
[0164] The present disclosure describes a model that adaptively reflects quantization error, which may include, first, a data-derived quantization model and, second, a simple equivalent quantization model.
[0165] According to one embodiment, the method according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0166] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device, At least one processor (120); and It includes a memory (130) that stores at least one instruction, and When the above at least one instruction is executed by at least part of the above at least one processor (120), it causes the electronic device to perform at least one operation, and The above at least one operation is: An operation to determine the maximum value of the output value of a first node included in a quantization-related artificial intelligence model and the maximum value; Operation to update the above maximum and minimum values; The operation of calculating a quantization unit based on the above-mentioned updated maximum and minimum values; and An electronic device characterized by including an operation to update quantization noise applied to the artificial intelligence model based on the quantization unit.
2. In paragraph 1, the above at least one operation is: The operation of performing training on an artificial intelligence model with the above-mentioned updated quantization noise applied; An operation to determine parameters for the above artificial intelligence model; and An electronic device characterized by including an operation of performing inference using the artificial intelligence model based on the parameters determined above.
3. An electronic device according to claim 1, characterized in that the output value includes the output value of a floating-point model included in the first node.
4. In paragraph 3, the operation of updating the quantization noise is, The operation of dividing the output values of the above floating-point model into the above quantization unit; The operation of calculating a first value for the output values of the above floating-point model by rounding to the nearest value among the values corresponding to the quantization unit, or by rounding down to the values corresponding to the quantization unit; and An electronic device characterized by further including the operation of multiplying the first value by the quantization unit.
5. In paragraph 1, the above at least one operation is, An electronic device characterized by including the operation of calculating the output value of the second node of the artificial intelligence model for the input value to which the above-mentioned updated quantization noise is applied.
6. In paragraph 1, the operation of updating the maximum and minimum values is, The operation of updating the maximum value includes the operation of updating the maximum value, and the operation of updating the maximum value The maximum value of the output of the first node is In this case, the first maximum value up to the j-th training data set of the first node ( ) is the second maximum value up to the j-1th training data set of the first node ( If it is greater than ), An electronic device characterized by including an operation to update to; 7. In a method for training an artificial intelligence model related to the quantization of electronic devices in a wireless communication system, An operation to determine the maximum value of the output value of a first node included in a quantization-related artificial intelligence model and the maximum value; Operation to update the above maximum and minimum values; The operation of calculating a quantization unit based on the above-mentioned updated maximum and minimum values; and A method characterized by including an operation to update quantization noise applied to the artificial intelligence model based on the quantization unit.
8. In Paragraph 7, The operation of performing training on an artificial intelligence model with the above-mentioned updated quantization noise applied; An operation to determine parameters for the above artificial intelligence model; and A method characterized by further including the operation of performing inference using the artificial intelligence model based on the determined parameters.
9. A method according to claim 7, wherein the output value includes the output value of a floating-point model included in the first node.
10. In claim 9, the operation of updating the quantization noise is, The operation of dividing the output values of the above floating-point model into the above quantization unit; The operation of calculating a first value for the output values of the above floating-point model by rounding to the nearest value among the values corresponding to the quantization unit, or by rounding down to the values corresponding to the quantization unit; and A method characterized by including the operation of multiplying the first value by the quantization unit.
11. In Paragraph 7, A method characterized by further including an operation to calculate the output value of the second node of the artificial intelligence model for the input value to which the above-mentioned updated quantization noise is applied.
12. In paragraph 7, the operation of updating the maximum and minimum values is, The operation of updating the maximum value includes the operation of updating the maximum value, and the operation of updating the maximum value The maximum value of the output of the first node is In this case, the first maximum value up to the j-th training data set of the first node ( ) is the second maximum value up to the j-1th training data set of the first node ( If it is greater than ), A method characterized by including an operation to update to 13. In paragraph 7, the operation of updating the maximum and minimum values is, The operation of updating the minimum value above is included, and the operation of updating the minimum value above is The minimum value of the output value of the first node is In this case, the first minimum value up to the j-th training data set of the first node ( ) is the second minimum value up to the j-1th training data set of the first node ( If it is smaller than ), A method characterized by including an operation to update to; 14. A storage medium storing at least one instruction readable by a computer, wherein the at least one instruction causes the electronic device to perform at least one operation when executed by at least a part of at least one processor (120) of the electronic device, and The above at least one operation is: An operation to determine the maximum value of the output value of a first node included in a quantization-related artificial intelligence model and the maximum value; Operation to update the above maximum and minimum values; The operation of calculating a quantization unit based on the above-mentioned updated maximum and minimum values; and A storage medium characterized by including an operation to update quantization noise applied to the artificial intelligence model based on the above quantization unit.
15. In paragraph 14, the above at least one operation is: The operation of performing training on an artificial intelligence model with the above-mentioned updated quantization noise applied; An operation to determine parameters for the above artificial intelligence model; and A storage medium characterized by further including the operation of performing inference using the artificial intelligence model based on the parameters determined above.
Citation Information
Patent Citations
Method for manufacturing aragonite from municipal waste and method for removing heavy metals in wastewater using the aragonite
KR1020240009866A
Dynamic quantization for deep neural network inference system and method
US20190012559A1
Dynamic quantization of neural networks
US20190042935A1
Vector clustered quantization
US20240354570A1
KR20240087565A