Electronic device for performing calculation using artificial intelligence model, and method for operating electronic device

By dynamically configuring AI models with optimized data types based on task requirements, the electronic device enhances performance metrics such as memory usage, computational speed, and power consumption while maintaining accuracy.

WO2025095481A1PCT designated stage expired Publication Date: 2025-05-08SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016440
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-10-25
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing AI models face challenges in optimizing memory usage, computational speed, and power consumption while maintaining accuracy, especially when using different data types for prompt executives and token generators.

Method used

The electronic device configures AI models with unique data types (FP32, FP16, INT8, INT4) learned using state-of-the-art architectures, allowing for dynamic data type selection based on task requirements to optimize performance metrics.

Benefits of technology

This approach enables the AI device to reduce memory usage, power consumption, and CPU usage while maintaining or improving accuracy, depending on the specific task and data type configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016440_08052025_PF_FP_ABST
    Figure KR2024016440_08052025_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device comprises a memory for storing a first artificial intelligence model, and a processor. The first artificial intelligence model includes: a prompt execution unit for executing a user prompt; and a token generation unit for generating a single token. The processor: measures at least one of the accuracy, a memory usage, or an operation speed of the first artificial intelligence model on the basis of the execution of instructions stored in the memory and stores the measured result in the memory; determines the type of first data to be provided to the prompt execution unit among the plurality of data types; determines the type of second data to be provided to the token generation unit; generates a second artificial intelligence model on the basis of the type of the first data and the type of the second data; and determines whether to use the second artificial intelligence model by comparing at least one of the accuracy, a memory usage, or an operation speed of the second artificial intelligence model with that of the first artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and operating methods for electronic devices that perform calculations using artificial intelligence models

[0001] This article relates to electronic devices, and more specifically, to electronic devices that perform computations using artificial intelligence models and methods of operating such devices.

[0002] Recent advancements in AI (Artificial Intelligence) technology and the increased computing power of mobile devices have led to AI technology being utilized in a variety of mobile applications. Countless devices, including smartphones, wearables, TVs, smart appliances, and AI speakers, can incorporate learning models to perform various functions, including device operation, measurement, outcome prediction, content recommendation, and decision-making. These models can be embedded and utilized early in device production or continuously improved through learning model updates from a server.

[0003] The above information may be provided as background information to aid in understanding this document. None of the above is claimed to be prior art related to this document or can be used to determine prior art.

[0004] If the AI ​​model type is FP32, the prompt execution unit and token generation unit can also use a single data type. The prompt execution unit can use the FP32 data type, and the token generation unit can also use the FP32 data type.

[0005] However, in actual examples, even if the prompt execution unit and token generation unit are determined to have different data types, the computational method can be identical to that of the original AI model. If the prompt execution unit and token generation unit are determined to have different data types from the original AI model, differences in data processing accuracy may occur, but this may vary depending on the type of task being processed (e.g., translation, response, typo correction, summary, etc.).

[0006] Compared to when the prompt execution unit and token generation unit uniformly use FP32 data types, using different data types can result in lower inference accuracy. However, the required inference accuracy may vary depending on the task type (e.g., translation, response, typo correction, summary, etc.). Therefore, because the required inference accuracy and performance vary depending on the task type, there may be situations where the prompt execution unit and token generation unit can sufficiently perform learning even when using different data types.

[0007] In one embodiment, models used in AI are trained using a state-of-the-art architecture as the basic design or are configured with unique data types (FP32, FP16, INT8, INT4) tailored to the specific purpose. Furthermore, the data types used during the actual model training phase are also used for the actual AI model inference process.

[0008] Among AI MODEL fields, the actual size of large language models (LLMs) is enormous, and various methods are currently being implemented to reduce this. For example, active research is underway to reduce the actual model size of language models like LLMs through 4-bit quantization. In fact, models with the 4-bit quantization of the LLM small model can be implemented on handset devices.

[0009] 4-bit quantization has a physical bit count that's 1 / 4 that of FP32, which reduces the model size. However, this inevitably increases the probability of drops (or omissions) in accuracy. This paper proposes a method that achieves results relatively similar to the original (FP32) model by modifying the model configuration according to data type, compared to using a single 4-bit quantization model.

[0010] An electronic device includes a memory storing a first artificial intelligence model, and a processor, wherein the first artificial intelligence model includes a prompt execution unit that executes a user prompt, and a token generation unit that generates a single token, and the processor measures at least one of accuracy, memory usage, or operation speed of the first artificial intelligence model based on execution of instructions stored in the memory and stores the measured results in the memory, determines a type of first data to be provided to the prompt execution unit among a plurality of data types, determines a type of second data to be provided to the token generation unit, generates a second artificial intelligence model based on the type of the first data and the type of the second data, and compares at least one of accuracy, memory usage, or operation speed of the second artificial intelligence model with the first artificial intelligence model to determine whether to use the second artificial intelligence model.

[0011] The operating method may include an operation of measuring at least one of accuracy, memory usage, or operation speed of a first artificial intelligence model and storing the measurement in memory, an operation of determining a type of first data to be provided to a prompt execution unit among a plurality of data types, an operation of determining a type of second data to be provided to a token generation unit, an operation of generating a second artificial intelligence model based on the type of the first data and the type of the second data, and an operation of comparing at least one of accuracy, memory usage, or operation speed of the generated second artificial intelligence model with the first artificial intelligence model to determine whether to use the second artificial intelligence model.

[0012] The recording medium may include a memory storing a first artificial intelligence model, and a processor, and the first artificial intelligence model may include a prompt execution unit that executes a user prompt, and a token generation unit that generates a single token. The memory may store instructions that, when executed, control the processor to measure at least one of accuracy, memory usage, or operation speed of the first artificial intelligence model and store the measured result in the memory, determine a type of first data to be provided to the prompt execution unit among a plurality of data types, determine a type of second data to be provided to the token generation unit, generate a second artificial intelligence model based on the type of the first data and the type of the second data, and compare at least one of accuracy, memory usage, or operation speed of the second artificial intelligence model with the first artificial intelligence model to determine whether to use the second artificial intelligence model.

[0013] Compared to when the prompt execution unit and the token generation unit all use the FP32 data type, using different data types can save memory capacity, reduce inference time, relatively reduce power consumption, and reduce CPU usage. Therefore, the electronic device according to this document can determine the accuracy required for each type of task and determine the first data type of the prompt execution unit and the second data type of the token generation unit differently to perform the task without errors while reducing memory usage, reducing power consumption, and improving the operation speed.

[0014] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0015] FIG. 2 is a block diagram illustrating an integrated intelligence system according to one embodiment.

[0016] FIG. 3 is a block diagram illustrating the configuration of an electronic device and a server according to various embodiments.

[0017] FIG. 4 is a block diagram illustrating a learning model management system according to various embodiments.

[0018] Figure 5a illustrates a token generation method in an artificial intelligence model according to a comparative example.

[0019] Figure 5b illustrates the configuration of an artificial intelligence model according to a comparative example.

[0020] Figure 5c is a diagram for explaining static allocation.

[0021] Figure 5d is a diagram for explaining dynamic allocation.

[0022] Figures 5e and 5f illustrate models generated with a single data type according to a comparative example.

[0023] FIG. 6 is a block diagram illustrating a method for configuring an artificial intelligence model by an electronic device according to various embodiments.

[0024] FIG. 7a and FIG. 7b illustrate a process of distinguishing models that satisfy performance among a plurality of artificial intelligence model combinations according to various embodiments.

[0025] FIG. 8A and FIG. 8B illustrate an example in which an electronic device according to various embodiments configures different artificial intelligence models according to performance for the same task.

[0026] FIG. 9 is a flowchart illustrating a method for an electronic device to configure an artificial intelligence model according to various embodiments.

[0027] FIG. 10 is a flowchart illustrating a method for an electronic device to configure an artificial intelligence model according to various embodiments.

[0028] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0029] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0030] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0031] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0032] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0033] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0034] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0035] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0036] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0037] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0038] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0039] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0040] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0041] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0042] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0043] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0044] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0045] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0046] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0047] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0048] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0049] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0050] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0051] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0052] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0053] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0054] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0055] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0056] FIG. 2 is a block diagram illustrating an integrated intelligence system according to various embodiments.

[0057] Referring to FIG. 2, according to one embodiment, the integrated intelligence system may include an electronic device (210) (e.g., electronic device (101) of FIG. 1), an intelligent server (230) (e.g., server (108) of FIG. 1), and a service server (250) (e.g., server (108) of FIG. 1).

[0058] According to one embodiment, the electronic device (210) may be a terminal device (or electronic device) that can connect to the Internet, for example, a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a TV, white goods, a wearable device, an HMD, or a smart speaker.

[0059] According to the illustrated embodiment, the electronic device (210) may include a communication interface (213) (e.g., the interface (177) of FIG. 1), a microphone (212) (e.g., the input module (150) of FIG. 1), a speaker (216) (e.g., the audio output module (155) of FIG. 1), a display module (211) (e.g., the display module (160) of FIG. 1), a memory (215) (e.g., the memory (130) of FIG. 1), or a processor (214) (e.g., the processor (120) of FIG. 1). The components listed above may be operatively or electrically connected to each other. The electronic device (210) may include at least some of the configurations and / or functions of the electronic device (101) of FIG. 1.

[0060] In one embodiment, the communication interface (213) may be configured to connect to an external device and transmit and receive data. In one embodiment, the microphone (212) may receive sound (e.g., user speech) and convert it into an electrical signal. In one embodiment, the speaker (216) may output the electrical signal as sound (e.g., voice).

[0061] In one embodiment, the display module (211) may be configured to display an image or video. In one embodiment, the display module (211) may also display a graphical user interface (GUI) of a running app (or application program). In one embodiment, the display module (211) may receive a touch input via a touch sensor. For example, the display module (211) may receive a text input via a touch sensor in an on-screen keyboard area displayed within the display module (211).

[0062] According to one embodiment, the memory (215) may store a client module (218), a software development kit (SDK) (217), and a plurality of apps (219a, 219b). The client module (218) and the SDK (217) may constitute a framework (or solution program) for performing general-purpose functions. In addition, the client module (218) or the SDK (217) may constitute a framework for processing user input (e.g., voice input, text input, touch input).

[0063] According to one embodiment, the plurality of apps (219a, 219b) stored in the memory (215) may be programs for performing a specified function. According to one embodiment, the plurality of apps may include a first app (219a) and a second app (219b). According to one embodiment, each of the plurality of apps (219a, 219b) may include a plurality of operations for performing a specified function. For example, the apps (219a, 219b) may include an alarm app, a message app, and / or a schedule app. According to one embodiment, the plurality of apps (219a, 219b) may be executed by the processor (214) to sequentially execute at least some of the plurality of operations.

[0064] According to one embodiment, the processor (214) can control the overall operation of the electronic device (210). For example, the processor (214) can be electrically connected to a communication interface (213), a microphone (212), a speaker (216), and a display module (211) to perform a designated operation.

[0065] According to one embodiment, the processor (214) may also execute a program stored in the memory (215) to perform a designated function. For example, the processor (214) may execute at least one of the client module (218) or the SDK (217) to perform the following operations for processing user input. The processor (214) may control the operations of a plurality of apps (219a, 219b), for example, through the SDK (217). The following operations described as operations of the client module (218) or the SDK (217) may be operations executed by the processor (214).

[0066] According to one embodiment, the client module (218) can receive user input. For example, the client module (218) can receive a voice signal corresponding to a user utterance detected through the microphone (212). Alternatively, the client module (218) can receive a touch input detected through the display module (211). Alternatively, the client module (218) can receive a text input detected through a keyboard or a visual keyboard. In addition, the client module (218) can receive various forms of user input detected through an input module included in the electronic device (210) or an input module connected to the electronic device (210). The client module (218) can transmit the received user input to the intelligent server (230). The client module (218) can transmit status information of the electronic device (210) together with the received user input to the intelligent server (230). The status information can be, for example, execution status information of an app.

[0067] In one embodiment, the client module (218) may receive a result corresponding to the received user input. For example, the client module (218) may receive a result corresponding to the received user input if the intelligent server (230) can produce a result corresponding to the received user input. The client module (218) may display the received result on the display module (211). Additionally, the client module (218) may output the received result as audio through the speaker (216).

[0068] According to one embodiment, the client module (218) can receive a plan corresponding to the received user input. The client module (218) can display the results of executing multiple operations of the app according to the plan on the display module (211). For example, the client module (218) can sequentially display the results of executing multiple operations on the display module (211) and output audio through the speaker (216). The electronic device (210) can, for another example, display only some results of executing multiple operations (e.g., the result of the last operation) on the display module (211) and output audio through the speaker (216).

[0069] In one embodiment, the client module (218) may receive a request from the intelligent server (230) to obtain information necessary to produce a result corresponding to the voice input. In one embodiment, the client module (218) may transmit the necessary information to the intelligent server (230) in response to the request.

[0070] According to one embodiment, the client module (218) may transmit result information of executing multiple operations according to a plan to the intelligent server (230). The intelligent server (230) may use the result information to confirm that the received user input has been processed correctly.

[0071] In one embodiment, the client module (218) may include a voice recognition module. In one embodiment, the client module (218) may recognize voice inputs that perform limited functions through the voice recognition module. For example, the client module (218) may execute an intelligent app that processes voice inputs to perform organic actions based on a specified input (e.g., "Wake up!").

[0072] According to one embodiment, the intelligent server (230) can receive information related to a user voice input from an electronic device (210) via a communication network. According to one embodiment, the intelligent server (230) can convert data related to the received voice input into text data. According to one embodiment, the intelligent server (230) can generate a plan for performing a task corresponding to the user voice input based on the text data.

[0073] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI ​​system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)), or a combination of the above or another AI system. In one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI ​​system may select at least one plan from a plurality of predefined plans.

[0074] According to one embodiment, the intelligent server (230) may transmit the results according to the generated plan to the electronic device (210), or transmit the generated plan to the electronic device (210). According to one embodiment, the electronic device (210) may display the results according to the plan on the display module (211). According to one embodiment, the electronic device (210) may display the results of executing an operation according to the plan on the display module (211).

[0075] According to one embodiment, the intelligent server (230) may include a front end (231), a natural language platform (232), a capsule database (238), an execution engine (233), an end user interface (234), a management platform (235), a big data platform (236), or an analytic platform (237).

[0076] According to one embodiment, the front end (231) can receive user input from the electronic device (210). The front end (231) can transmit a response corresponding to the user input.

[0077] According to one embodiment, the natural language platform (232) may include an automatic speech recognition module (ASR module) (232a), a natural language understanding module (NLU module) (232b), a planner module (232c), a natural language generator module (NLG module) (232d), or a text to speech module (TTS module) (232e).

[0078] According to one embodiment, the automatic speech recognition module (232a) can convert voice input received from the electronic device (210) into text data. According to one embodiment, the natural language understanding module (232b) can use the text data of the voice input to determine the user's intent. For example, the natural language understanding module (232b) can perform syntactic analysis or semantic analysis on user input in the form of text data to determine the user's intent. According to one embodiment, the natural language understanding module (232b) can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from the voice input, and can match the meaning of the determined words to the intent to determine the user's intent. The natural language understanding module (223b) can obtain intent information corresponding to the user's utterance. The intent information can be information indicating the user's intent determined by interpreting text data. The intent information can include information indicating an action or function that the user intends to execute using the device.

[0079] According to one embodiment, the planner module (232c) can generate a plan using the intent and parameters determined by the natural language understanding module (232b). According to one embodiment, the planner module (232c) can determine a plurality of domains necessary to perform a task based on the determined intent. The planner module (232c) can determine a plurality of operations included in each of the plurality of domains determined based on the intent. According to one embodiment, the planner module (232c) can determine parameters necessary to execute the determined plurality of operations or result values ​​output by the execution of the plurality of operations. The parameters and the result values ​​can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and a plurality of concepts determined by the user's intent. The planner module (232c) can determine the relationships between the plurality of operations and the plurality of concepts in a stepwise (or hierarchical) manner. For example, the planner module (232c) can determine the execution order of a plurality of actions based on the user's intention based on a plurality of concepts. In other words, the planner module (232c) can determine the execution order of a plurality of actions based on parameters required for the execution of the plurality of actions and results output by the execution of the plurality of actions. Accordingly, the planner module (232c) can generate a plan including association information (e.g., ontology) between the plurality of actions and the plurality of concepts. The planner module (232c) can generate the plan using information stored in a capsule database that stores a set of relationships between concepts and actions.

[0080] According to one embodiment, the natural language generation module (232d) can convert specified information into text format. The information converted into text format may be in the form of natural language speech. According to one embodiment, the text-to-speech module (232e) can convert text-to-speech information into speech information.

[0081] According to one embodiment, some or all of the functions of the natural language platform (232) may also be implemented in the electronic device (210).

[0082] The capsule database can store information about the relationships between multiple concepts and actions corresponding to multiple domains. According to one embodiment, the capsule can include multiple action objects (or action information) and concept objects (or concept information) included in the plan. According to one embodiment, the capsule database can store multiple capsules in the form of a concept action network (CAN). According to one embodiment, the multiple capsules can be stored in a function registry included in the capsule database.

[0083] The capsule database may include a strategy registry that stores strategy information necessary for determining a plan corresponding to a user input. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to the user input. According to one embodiment, the capsule database may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a given situation. The follow-up actions may include, for example, follow-up utterances. According to one embodiment, the capsule database may include a layout registry that stores layout information of information output through the electronic device (210). According to one embodiment, the capsule database may include a vocabulary registry that stores vocabulary information included in the capsule information. According to one embodiment, the capsule database may include a dialog registry that stores information on dialogue (or interaction) with the user. The capsule database may update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies that determine plans. The developer tool may include a dialog editor for creating conversations with users.The developer tool may include a follow-up editor that activates follow-up goals and allows editing of follow-up utterances that provide hints. The follow-up goals may be determined based on the currently set goals, user preferences, or environmental conditions. In one embodiment, the capsule database may also be implemented within the electronic device (210).

[0084] In one embodiment, the execution engine (233) can use the generated plan to produce a result. The end user interface (234) can transmit the produced result to the electronic device (210). Accordingly, the electronic device (210) can receive the result and provide the received result to the user. In one embodiment, the management platform (235) can manage information used in the intelligent server (230). In one embodiment, the big data platform (236) can collect user data. In one embodiment, the analysis platform (237) can manage the quality of service (QoS) of the intelligent server (230). For example, the analysis platform (237) can manage the components and processing speed (or efficiency) of the intelligent server (230).

[0085] According to one embodiment, the service server (250) can provide a service (e.g., food ordering or hotel reservation) specified to the electronic device (210). According to one embodiment, the service server (250) can be a server operated by a third party. According to one embodiment, the service server (250) can provide information for generating a plan corresponding to the received voice input to the intelligent server (230). The provided information can be stored in a capsule database. In addition, the service server (250) can provide result information according to the plan to the intelligent server (230). The service server (250) can include a plurality of service providers (e.g., CP Service A (251), CP Service B (252), CP Service C (253)), and each of the service providers (251, 252, 253) can provide a function for a domain associated with each capsule stored in the capsule database (238) of the intelligent server (230).

[0086] In the integrated intelligence system described above, the electronic device (210) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.

[0087] According to one embodiment, the electronic device (210) may provide a voice recognition service through an intelligent app (or voice recognition app) stored within the device. In this case, for example, the electronic device (210) may recognize a user utterance or voice input received through the microphone (212) and provide the user with a service corresponding to the recognized voice input.

[0088] According to one embodiment, the electronic device (210) may perform a designated operation based on the received voice input, either alone or together with the intelligent server (230) and / or the service server (250). For example, the electronic device (210) may execute an app corresponding to the received voice input and perform a designated operation through the executed app.

[0089] According to one embodiment, when an electronic device (210) provides a service together with an intelligent server (230) and / or a service server (250), the electronic device (210) may detect a user's speech using the microphone (212) and generate a signal (or voice data) corresponding to the detected user's speech. The electronic device (210) may transmit the voice data to the intelligent server (230) via a network (240) using a communication interface (213).

[0090] In one embodiment, an intelligent server (230) may generate a plan for performing a task corresponding to a voice input received from an electronic device (210), or a result of performing an operation according to the plan, in response to the voice input. The plan may include, for example, a plurality of operations for performing a task corresponding to a user's voice input, and a plurality of concepts related to the plurality of operations. The concept may define parameters input to the execution of the plurality of operations, or result values ​​output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and the plurality of concepts.

[0091] According to one embodiment, the electronic device (210) can receive the response using the communication interface (213). The electronic device (210) can output a voice signal generated within the electronic device (210) to the outside using the speaker (216), or can output an image generated within the electronic device (210) to the outside using the display module (211).

[0092] Although FIG. 2 describes an example in which voice recognition of user input received from an electronic device (210), natural language understanding and generation, and calculation of results using a plan are performed on an intelligent server (230), the various embodiments of the present document are not limited thereto. For example, at least some components of the intelligent server (230) (e.g., natural language platform (232), execution engine (233), capsule database (238)) may be embedded in the electronic device (210) (or the electronic device (101) of FIG. 1), so that the operations may be performed by the electronic device (210).

[0093] FIG. 3 is a block diagram illustrating the configuration of an electronic device and a server according to various embodiments.

[0094] According to various embodiments, the electronic device (300) may include a processor (310), a communication module (320), and a first model (301-1), and some of the illustrated configurations may be omitted or replaced. The electronic device may further include at least some of the configurations and / or functions of the electronic device (101) of FIG. 1. At least some of the respective configurations of the illustrated (or not illustrated) electronic device may be operatively, functionally, and / or electrically connected to each other.

[0095] According to various embodiments, the processor (310) may be configured as one or more processors capable of performing calculations or data processing related to control and / or communication of each component of an electronic device. The processor (310) may include at least some of the configurations and / or functions of the processor (120) of FIG. 1.

[0096] According to various embodiments, there may be no limitation to the computational and data processing functions that the processor (310) may implement on the electronic device (300), but below, features related to the control of the learning model will be described in detail. The operations of the processor (310) may be performed by loading instructions stored in a memory (not shown).

[0097] According to various embodiments, the communication module (320) may communicate with an external device via a wireless network under the control of the processor (310). The communication module (320) may include hardware and software modules for transmitting and receiving data from a cellular network (e.g., a long term evolution (LTE) network, a 5G network, a new radio (NR) network) and a short-range network (e.g., Wi-Fi, Bluetooth). The communication module (320) may include at least some of the configurations and / or functions of the communication module (190) of FIG. 1.

[0098] The electronic device (300) according to one embodiment may be implemented in various forms. For example, the electronic device (300) described in this specification may include, but is not limited to, a smart TV, a set-top box, a mobile phone, a tablet PC, a digital camera, a laptop computer, a desktop, an e-book reader, a digital broadcasting terminal, a PDA (Personal Digital Assistant), a PMP (Portable Multimedia Player), a navigation device, an MP3 player, a wearable device, etc.

[0099] According to one embodiment, input information may refer to information input to the first model (301-1) to perform an operation. In addition, output information may refer to a result of processing the input information on the first model (301-1). The server may obtain output information of the first model (301-1). The server may perform various operations based on the output information of the first model (301-1). For example, the server may determine that the performance of an existing learning model installed on the terminal has deteriorated based on the output information of the first model (301-1). Based on the determination that the performance of the existing learning model has deteriorated, the server may learn a new learning model and distribute it to the terminal.

[0100] According to an embodiment, the first model (301-1) may be at least one artificial intelligence model for performing various operations in the electronic device (300). For example, the first model (301-1) may be an artificial intelligence model for various operations that may be performed in the electronic device (300), such as a voice recognition model, a natural language processing model, an image recognition model, etc. Without being limited to the above-described example, the first model (301-1) may be one of various types of artificial intelligence models.

[0101] According to one embodiment, the server (305) may transmit information related to updating the first model (301-1) to the communication module (320) on the electronic device (300) using the second model (303) that has better performance than the first model (301-2). The electronic device (300) may update the learning model using the information related to updating the first model (301-1). According to one embodiment, the electronic device (300) may continuously update the learning model based on the information provided from the server (305). The electronic device (300) may output output information corresponding to the input information using the updated learning model (e.g., the second model (303)).

[0102] In one embodiment, the entity performing training for a new learning model and the entity distributing the new learning model may be the same or different. For example, the server (305) may perform training for a new learning model and transmit or distribute the new learning model to the electronic device (300) based on receiving a signal requesting the new learning model from the electronic device (300). Alternatively, the server (305) may perform training for a new learning model and transmit the new learning model to another server (not shown). The other server (not shown) that receives the new learning model may transmit or distribute the new learning model to the electronic device (300) based on a request from the electronic device (300).

[0103] According to one embodiment, the second model (303) may have a relatively larger size than the first model (301-2), and thus may have a relatively higher probability of outputting information suitable for the user. According to one embodiment, the second model (303) may be an artificial intelligence model that includes a greater number of nodes and neural network layers than the first model (301-2), as it may be processed by a server (305) with relatively high performance.

[0104] According to one embodiment, the electronic device (300) may receive information indicating improvements to a new learning model from the model evaluation unit (330) of the server (305) using the communication module (302). The electronic device (300) may determine the need for a new learning model based on the information indicating improvements to the new learning model. In response to determining that an update to a new model is necessary, the electronic device (300) may request information about the new learning model from the server (305). In response to the request of the electronic device (300), the server (305) may transmit information about the new learning model (e.g., the second model (303)) to the electronic device (300). The electronic device (300) may update the learning model based on the information received from the server (305).

[0105] According to one embodiment, the server (305) can use the model evaluation unit (330) to identify differences or improvements between the first model (301-2) and the second model (303). The first model (301-2) may refer to a learning model installed in the electronic device (300) (the same model as the first model (301-1).

[0106] According to one embodiment, the server (305) may update the first model (301-2) installed in the server (305) using the second model (303), as an artificial intelligence model, before the electronic device (300). According to one embodiment, the server (305) may update the first model (301-2) so that the same output information as the output information output from the second model (303) can be output for the same input information. In addition, the server (305) may obtain information for updating the first model (301-1) of the electronic device (300) based on the updated first model (301-2) and transmit the information to the electronic device (300). Alternatively, the electronic device (300) may transmit information indicating differences and improvements between the first model (301-2) and the second model (303) to the electronic device (300).

[0107] In one embodiment, improvements may occur when the data distribution input to the learning model changes. Alternatively, improvements may occur when a specific target class that is the subject of analysis for the learning model is improved. Based on improvements in the specific target class that is the subject of analysis for the learning model, the server (305) may transmit the improved learning model only to terminals that require improvement in the corresponding class. The server (305) may not transmit the learning model to terminals that do not require improvement in the corresponding class, thereby saving resources for the server (305) and the corresponding terminals.

[0108] FIG. 4 is a block diagram illustrating a learning model management system according to various embodiments.

[0109] The server (405) may include the server (305) of the preceding FIG. 3. The electronic device (400) may include the electronic device (300) of the preceding FIG. 3.

[0110] In operation (1), a model evaluator (415) of a server (e.g., server (305) of FIG. 3) (405) collects the current model performance of an electronic device (e.g., electronic device (300) of FIG. 3) (400), and based on the verification result of the model performance, the electronic device (400) can determine whether a new learning model is required.

[0111] In operation (2), the model evaluator (415) can control the trainer (435) to perform new learning using the model manager (425) based on determining that a new learning model is needed.

[0112] In operation (3), the trainer (435) can train a new trained model (445). The feature analyzer (455) can analyze the new trained model (445) and improvements of the existing trained model and store them in the model manager (425).

[0113] In operation (4), the model manager (425) can compare the existing learning model of the electronic device (400) with the new learning model and transmit the improvements to the electronic device (400). The model manager (410) of the electronic device (400) can determine whether the data distribution of the new learning model is similar to the data distribution of the existing learning model using the input data distribution evaluator (411). The model manager (410) of the electronic device (400) can determine whether a specific target class of the new learning model has been improved and whether improvement of the target class is required on the electronic device (400) using the target class evaluator (412). According to one embodiment, the electronic device (400) can analyze the user's usage pattern and determine that improvement of the target class is required based on the use of the improved target class at a certain frequency or more.

[0114] In one embodiment, the electronic device (400) may transmit information indicating that a new learning model (445) is needed to the server (405) based on determining that a new learning model (445) is needed.

[0115] In operation (5), the processor (420) of the electronic device (400) can download a new learning model (445) from the server (405) and update the learning model within the electronic device (400). According to a comparative embodiment, the electronic device may waste resources by downloading the new learning model (445) even when there is no need to update the learning model. According to various embodiments, the electronic device (400) can download only the improvements of the new learning model, not the new learning model, from the server (405), determine whether to update the learning model, and download the new learning model (445) so as not to waste resources.

[0116] In operation (6), the electronic device (400) can perform a prediction using a learning model using a prediction analyzer (430) and transmit feedback data to a model manager (410) of the electronic device (400). The model manager (410) of the electronic device (400) can determine an update time of the learning model based on the received feedback data.

[0117] Figure 5a illustrates a token generation method in an artificial intelligence model according to a comparative example.

[0118] Figure 5a illustrates a method of generating tokens in an artificial intelligence model (e.g., a large language model (LLM)) having a sequence to sequence structure. A sequence may refer to a series of data with an order. The sequence to sequence method may refer to a method of inputting a sequence that has undergone natural language processing (NLP) and outputting another sequence.

[0119] An AI model can process user input from an input prompt all at once. The AI ​​model can then iterate through inference, generating a single token at each iteration.

[0120] For example, in Figure 502, the AI ​​model can be input with tokens corresponding to recite, the, first, and law. While performing inference, the AI ​​model can generate a new token indicating A in Figure 504. Afterwards, the AI ​​model can generate a new token indicating 'robot' in Figure 506, and a new token indicating 'may' in Figure 508. The AI ​​model can generate a new token indicating 'not' in Figure 510. The AI ​​model can repeat inference until an EOS (end of sequence) token indicating the end of inference is generated.

[0121] In Figure 5a, the AI ​​model can repeatedly perform inference while retaining the tokens corresponding to "recite," "the," "first," and "law" in memory in their original sizes. The AI ​​model can continue to retain the tokens corresponding to "recite," "the," "first," and "law" in memory in their original sizes while performing inference-related operations 504 to 510. The AI ​​model can iterate through learning to adjust weights and reduce prediction errors.

[0122] AI models can generate a single token through a softmax operation while iterating through learning. To generate a single token, AI models can utilize a buffer, which is allocated temporary storage space. The softmax operation can refer to an operation that allows output values ​​to be interpreted as probabilities. In classification problems, it is used as an activation function in the output layer and can be used to indicate the probability of belonging to each class. A token can refer to the smallest unit of text for processing. A single token can refer to one token among multiple tokens. A buffer can refer to an area for temporarily storing data.

[0123] The AI ​​model can use additional memory equal to the LLM model's max_sequence_length and token_embedding_size . max_sequence_length can refer to the maximum length of tokens the model can process. token_embedding_size can refer to one of the vectors representing the tokens.

[0124] The values ​​generated from each iteration can be related to the results of previous operations. Because of this relationship, AI models may have difficulty performing subsequent operations without data on the results of previous operations. Therefore, AI models can store both the values ​​of the input prompts and the results of previous iterations in memory.

[0125] The values ​​generated by an AI model can be related to the results of previous computations. Therefore, an AI model can store both input prompts and the results of previously performed inferences in memory.

[0126] Figures 5b and 5c illustrate the configuration of an artificial intelligence model according to a comparative example.

[0127] In FIG. 5b, the artificial intelligence model (e.g., LLM) (520) may be configured with a prompt execution unit (522) and a token generation unit (524) to use a sequence-to-seq method. The prompt execution unit (522) may include, for example, a BERT (bidirectional encoder representations from transformers) model as a model for processing an input prompt. The token generation unit (524) may include a KVcache (key-value cache). A key-value cache may refer to a cache system that stores intermediate results to perform a specific task relatively quickly. The token generation unit (524) stores a key and a corresponding value as data, and can quickly find a value using the key.

[0128] In one embodiment, BERT can process words bidirectionally to simultaneously consider all words in a sentence. Using a model called a transformer, BERT can focus on important words and de-emphasize less important ones. BERT can be pre-trained using large amounts of text data. Through pre-training, BERT can acquire basic language patterns. BERT can be further fine-tuned for specialized or segmented tasks. Labeled data can be utilized for this fine-tuning.

[0129] In one embodiment, the prompt execution unit (522) can only process input prompts. Except for tokens generated through inference and data required for iteration (e.g., tokens corresponding to "recite," "the," "first," and "law" in FIG. 5A ), the remaining data can be loaded into memory.

[0130] According to one embodiment, the token generation unit (524) can receive a token generated from a previous inference as input and generate the next token.

[0131] Figure 5c is a diagram for explaining static allocation.

[0132] In one embodiment, in the case of static allocation, if the sequence length is 8 steps and the token size is 10 space (vector), the AI ​​model may require 8 * 10 = 80 memory capacity for each operation to operate. The AI ​​model may additionally use memory equal to the max_sequence_length and token_embedding_size of the LLM model. max_sequence_length may mean the maximum length of a token that the model can process. token_embedding_size may mean one of the vectors representing the token. For example, if max_sequence_length is 8 steps and token_embedding_size is a space (vector) with a size of 10, the model may always require 8 * 10 = 80 allocation to operate.

[0133] That is, even in a situation where repetitive learning is taking place, the artificial intelligence model must continue to use data of size 80, so it may require a relatively large memory capacity compared to dynamic allocation.

[0134] Figure 5d is a diagram for explaining dynamic allocation.

[0135] In one embodiment, the dynamic allocation method, on the other hand, has a structure in which only the token size of 10 is allocated and added to the existing information each time a new token is created by performing inference. Therefore, compared to the static allocation method, which requires 80 units of capacity for each operation, the dynamic allocation method can perform learning and inference using relatively small memory capacity.

[0136] In the case of dynamic allocation, only 10 token_embedding_size is allocated for each iteration of the learning process and can have a structure in which it is added to the existing information. According to Fig. 5d, the AI ​​model can allocate only 10 tokens of data in the first learning process. In the second learning process, the AI ​​model can add 10 new data for the second column while storing the existing 10 data in the memory. In the case of static allocation, the AI ​​model can perform learning using data with a size of 80 (=8*10) for each learning process. In the case of dynamic allocation, the AI ​​model learns by adding only 10 data for each learning process, so it can perform learning and inference using relatively less memory capacity than static allocation.

[0137] AI models can perform inference using relatively less memory than static allocation methods if the end of sequence (EOS) is generated before max_sequence_length. max_sequence_length can indicate the maximum length of tokens the model can process. EOS (end of sequence) can indicate the end of the operation. AI models can use EOS to determine where the input sequence ends.

[0138] Figures 5e and 5f illustrate models generated using a single data type according to comparative examples. In Figure 552, the AI ​​model is of the FP32 type and can have a size of 32 gigabytes. Since the AI ​​model type is FP32, the prompt execution unit (e.g., BERT) and token generation unit (e.g., KVACHE) can also use a single data type. The prompt execution unit uses the FP32 data type, and the token generation unit can also use the FP32 data type.

[0139] In Figure 554, since the type of the artificial intelligence model is FP16, it can be confirmed that the prompt execution part uses the data type of FP16, and the token generation part also uses the data type of FP16. According to one embodiment, a single model created with one data type, such as 64 bits (FP64, INT64), 32 bits (FP32, INT32), 16 bits (FP16, INT16), 8 bits (FP8, INT8), 4 bits (FP4, INT4), 2 bits (FP2, INT2), a model that adds a Bert model or a Kvcache model can be used. According to one embodiment, it can be composed of only a combination of models having the same data type, such as FP32, FP16, INT8, and INT4, and stored in the file system.

[0140] In Figure 556, since the type of the artificial intelligence model is INT8, it can be confirmed that the prompt execution part uses the data type of INT8, and the token generation part also uses the data type of INT8.

[0141] In Figure 554, since the type of the artificial intelligence model is INT4, it can be confirmed that the prompt execution part uses the data type of INT4, and the token generation part also uses the data type of INT4.

[0142] However, in actual examples, even if the prompt execution unit and token generation unit are determined to have different data types, the computational method can be identical to that of the original AI model. If the prompt execution unit and token generation unit are determined to have different data types from the original AI model, differences in data processing accuracy may occur, but this may vary depending on the type of task being processed (e.g., translation, response, typo correction, summary).

[0143] If the original AI model's data type is FP32, the prompt execution unit and token generation unit according to the comparative example use the same FP32 as the original. However, in the present invention, a different data type may be used. For example, even if the original uses FP32, the prompt execution unit of the present invention's AI model may use the FP16 data type, and the token generation unit may use the INT 8 data type.

[0144] When the prompt execution unit and the token generation unit use different data types compared to when they all use the FP32 data type, memory capacity can be saved, inference time can be reduced, power consumption can be relatively reduced, and CPU usage can be reduced. When the prompt execution unit and the token generation unit use different data types compared to when they all use the FP32 data type, there is a disadvantage that inference accuracy may be reduced. However, the required inference accuracy may vary depending on the type of task (e.g., translation, response, typo correction, summary, etc.). Therefore, the electronic device according to the present document determines the required accuracy depending on the type of task, and determines the first data type of the prompt execution unit and the second data type of the token generation unit differently, thereby performing the task without errors while reducing memory usage, reducing power consumption, and improving the operation speed. Hereinafter, the first data type may refer to the data type of the prompt execution unit. The second data type may refer to the data type of the token generation unit.

[0145] According to Fig. 5f, the artificial intelligence model can determine how to display data by selecting either the prompt execution part (e.g., BERT model) or the token generation part (e.g., KVcache model).

[0146] AI models can have a structure that uses pre-determined data types in the prompt execution unit (e.g., the Bert model) and token generation unit (e.g., the Kvcache model). Even when quantized into different data types, AI models can maintain the same computational methods as the original (FP32) model, with only differences in accuracy. Quantization can be used to reduce the size of AI models and accelerate computation. For example, an electronic device can convert a model from 32-bit floating point to 8-bit integers, thereby reducing the size of the AI ​​model, reducing memory usage, and improving computational speed.

[0147] Also, 32bit / 16bit / 8bit / 4bit / 2bit quantization may show relatively better performance than the original (64bit, 32bit) model compared to 64bit / 32bit. Therefore, even when using the same original model, an AI model with quantization can be used. An electronic device can configure an AI model with a combination of data types optimized for each different task (e.g., language translation, summarization, typo correction, response) to optimize at least one of inference time, power consumption, ROM size, RAM size, RAM transaction, backend utilization, or CPU occupation. Optimization may mean an operation that maximizes the performance of at least one of inference time, power consumption, ROM size, RAM size, RAM transaction, backend utilization, or CPU occupation.

[0148] FIG. 6 is a block diagram illustrating a method for configuring an artificial intelligence model by an electronic device according to various embodiments.

[0149] In Figure 6, Figure 610 illustrates the first data type of the prompt execution unit in block diagram form. Figure 630 illustrates the second data type of the token generation unit in block diagram form.

[0150] Figure 620 illustrates, in blocks, the process of selecting a data type for a prompt execution unit and a data type for a token generation unit within an artificial intelligence model. The model generation in Figure 620 can be executed during build time or runtime. Build time may refer to the time it takes for an artificial intelligence model to be built. The electronic device (300) may define the structure of the artificial intelligence model during build time and train the artificial intelligence model using training data. Runtime may refer to the time it takes for the artificial intelligence model to be executed and produce a result. After the artificial intelligence model is built through build time, it may perform predictions based on input data.

[0151] The operation of selecting a data type of the prompt execution unit and selecting a data type of the token generation unit can be performed by a processor (e.g., processor (310) of FIG. 3).

[0152] According to one embodiment, an electronic device (e.g., the electronic device (300) of FIG. 3) may select a model having a specific data type at build-time in a situation of performing specific tasks (e.g., language translation, summarization, typo correction, response). The electronic device (300) may perform specific tasks by selectively storing only specific models in ROM. Alternatively, the electronic device (300) may store an artificial intelligence model having multiple data types in ROM and configure a model composition for performing specific tasks by selecting any one of the multiple data types stored in the ROM at runtime.

[0153] The processor (310) or the electronic device (300) may select a first data type of the prompt execution unit and a second data type of the token generation unit when performing a specific task, thereby configuring a new artificial intelligence model. If there are pre-established conditions (rules), the electronic device (300) may determine the types of the first data and the second data based on the established conditions, and configure the artificial intelligence model according to the determined data types. Once the data types are determined, the artificial intelligence model may perform inference on the input data to generate or output a new token.

[0154] In one embodiment, the data types may include FP32, FP16, INT8, INT4, and INT2. The job or task may include any one of language translator, summarization, typo correction, or reply.

[0155] FIG. 7a and FIG. 7b illustrate a process of distinguishing models that satisfy performance among a plurality of artificial intelligence model combinations according to various embodiments.

[0156] According to FIG. 7a, an electronic device (e.g., the electronic device (300) of FIG. 3) can construct an artificial intelligence model using various data types, including an original artificial intelligence model. The electronic device (300) can determine the data type of a new artificial intelligence model based on at least one of the size of a read-only memory (ROM), the size of a random access memory (RAM), or a required computational speed (KPI, token / s). The electronic device (300) can determine the data types of a prompt execution unit and a token generation unit.

[0157] The electronic device (300) can configure an AI model based on a pre-established guide for model composition if one exists. In situations where no guide exists for model composition, the electronic device (300) can configure an AI model by considering specific components. Specific components may include, for example, RAM size, backend capability, or a key performance indicator (KPI). The KPI may refer to the number of tokens per second (tokens / sec) processed by the BERT model or the Kvcache model. For example, the more tokens the BERT model processes per second, the faster the AI ​​model's processing speed. A faster AI model's processing speed can improve the user experience and has the advantage of being able to process a relatively larger dataset.

[0158] For example, in Figure 710, the electronic device (300) can determine the data type so that the ROM size is 32 gigabytes, the RAM size is 8 gigabytes, and the operation speed satisfies 12 tokens / s. The ROM size, the RAM size, and the operation speed are only examples, and the conditions that the electronic device (300) considers to determine the data type are not limited thereto. The electronic device (300) can consider the performance of each combination of data. An indicator representing the performance of each combination of data may include, for example, at least one of the size of the memory, the memory capacity (capability) of the backend, the operation speed (key performance indicators, KPI), power consumption, or the accuracy of inference. In Figure 710, the electronic device (300) can create a model pool by collecting all data types that satisfy 32 gigabytes of ROM, 8 gigabytes of RAM, and 12 tokens / s. The model pool will be described in Figure 7b.

[0159] In Figure 720, the electronic device (300) can determine the data type to satisfy the following conditions: ROM size of 128 gigabytes, RAM size of 12 gigabytes, and operation speed of 18 token / s.

[0160] In Figure 730, the electronic device (300) can determine the data type to satisfy the following conditions: ROM size of 64 gigabytes, RAM size of 12 gigabytes, and operation speed of 16 token / s.

[0161] In Figure 740, the electronic device (300) can determine the data type to satisfy the size of ROM of 256 gigabytes, the size of RAM of 128 gigabytes, and the operation speed of 20 token / s.

[0162] The figures and elements mentioned in Figures 710 to 740 are merely examples. The figures used to determine the data types used in an AI model may vary depending on the type of task being performed. The elements used to determine the data types used in an AI model may vary depending on user settings.

[0163] Figure 7b is a block diagram showing combinations of data types that satisfy the conditions required to determine a data type.

[0164] For example, in Figure 710, the electronic device (300) can determine the data type so that the ROM size is 32 gigabytes, the RAM size is 8 gigabytes, and the operation speed is 12 tokens / s. The combination of data types that satisfy the conditions of Figure 710 may include, for example, a case where the prompt execution unit (e.g., BERT model) has a data type of INT4 and a token generation unit (e.g., KVcache model) has a data type of INT8. Alternatively, the conditions of Figure 710 may be satisfied when the prompt execution unit has a data type of either INT2 or INT4 and the token generation unit has a data type of either INT2, INT4, or INT8.

[0165] In one embodiment, the prompt execution unit may have a data type of INT4 (4 bits), a size of 4 gigabytes (G), and a specification to process 20 tokens per second. The token generation unit may have a data type of INT4 (4 bits), a size of 2 gigabytes (G), and a specification to process 20 tokens per second. If the prompt execution unit and the token generation unit satisfy these specifications, the condition of Figure 710 may be satisfied. This is merely an example, and the combination of the prompt execution unit and the token generation unit that satisfy the condition of Figure 710 is not limited to this.

[0166] The electronic device (300) can check the specifications of the prompt execution unit and the token generation unit, and generate a model composition based on the required conditions (e.g., the conditions of Figure 710). The electronic device (300) can determine a candidate group of the prompt execution unit and the token generation unit that satisfy the specifications based on at least one criterion among KPI, memory capacity, power consumption, and inference accuracy. The electronic device (300) can also perform auto composition that automatically selects one of the candidate groups of the prompt execution unit and the token generation unit that satisfy the specifications. For example, two types of prompt execution units that satisfy the specifications of Figure 710 are illustrated. Three types of token generation units that satisfy the specifications of Figure 710 are illustrated. The electronic device (300) can select one of the two types of prompt execution units and one of the three types of token generation units to perform auto composition. According to one embodiment, the electronic device (300) can determine the specifications of the artificial intelligence model including Figures 710 to 740 depending on the type of task to be performed. The specifications of the artificial intelligence model illustrated in Fig. 7b are merely examples and are not limited to those illustrated in Figs. 710 to 740. Depending on the specifications of the artificial intelligence model, the data type of the prompt execution unit and the data type of the token generation unit may be determined differently. The electronic device (300) determines the specifications of the artificial intelligence model depending on the type of task being performed, and depending on the specifications of the artificial intelligence model, the data type of the prompt execution unit and the data type of the token generation unit may be determined differently.

[0167] In one embodiment, as the specifications of an AI model improve, the types of data types that can be combined may also increase. For example, in the case of Figure 710, the ROM size is 32 gigabytes, the RAM size is 8 gigabytes, and the operation speed satisfies 12 tokens / s, so the specifications may be relatively lower than in Figure 740. In the case of Figure 710, the prompt execution unit can have either INT2 or INT4 data types, and the token generation unit can have either INT2, INT4, or INT8 data types.

[0168] On the other hand, in the case of Figure 740, the specifications of the allowed artificial intelligence model are relatively good, so data types can be combined relatively more than in Figure 710. For example, in the case of Figure 740, the prompt execution unit can have a data type of any one of INT2, INT4, INT8, FP16, or FP32, and the token generation unit can have a data type of any one of INT2, INT4, INT8, FP16, FP32, or FP64.

[0169] FIG. 8A and FIG. 8B illustrate an example in which an electronic device according to various embodiments configures different artificial intelligence models according to performance for the same task.

[0170] FIG. 8a illustrates an embodiment in which the configuration of an artificial intelligence model is varied when processing the same type of task on different electronic devices (810, 820, and 830).

[0171] In FIG. 8A, the first electronic device (810) may have a data type of FP32 or FP16. In FIG. 812, when processing the first task, the first electronic device (810) may determine the first data type of the prompt execution unit as FP32, and the second data type of the token generation unit as FP32. Among the artificial intelligence models, the data type of the conventional model may be set to be the same regardless of the performance of the electronic device (300). Among the artificial intelligence models, the data type of the dynamic model may be set differently depending on the performance of the electronic device (300).

[0172] In Figure 814, when the first electronic device (810) processes the second task, the first data type of the prompt execution unit can be determined as FP16, and the second data type of the token generation unit can be determined as FP32.

[0173] In Figure 816, when the first electronic device (810) processes the third task, the first data type of the prompt execution unit may be determined as FP32, and the second data type of the token generation unit may be determined as FP16. This is merely an example, and the first data type of the prompt execution unit and the second data type of the token generation unit may vary depending on the specifications of the first electronic device (810) and the specifications required by the task being performed.

[0174] In FIG. 8A, the second electronic device (820) may have a data type of FP32, INT8, or INT4. In this case, in FIG. 822, when the second electronic device (820) processes the first task, the first data type of the prompt execution unit may be determined as FP32, and the second data type of the token generation unit may be determined as INT4.

[0175] In Figure 824, when the second electronic device (820) processes the second task, the first data type of the prompt execution unit can be determined as INT8, and the second data type of the token generation unit can be determined as INT8.

[0176] In Figure 826, the second electronic device (820) may determine the first data type of the prompt execution unit as INT8 and the second data type of the token generation unit as INT4 when processing the third task. This is merely an example, and the first data type of the prompt execution unit and the second data type of the token generation unit may vary depending on the specifications of the second electronic device (820) and the specifications required by the task being performed.

[0177] In FIG. 8A, the third electronic device (830) may have a data type of INT8, INT4, or INT2. In this case, in FIG. 832, when the third electronic device (830) processes the first task, the first data type of the prompt execution unit may be determined as INT8, and the second data type of the token generation unit may be determined as INT2.

[0178] In FIG. 8A, the method of model composition for performing a task may vary depending on the electronic device (e.g., terminal, computer). For example, a first electronic device (A device) (810) may configure an AI model for performing a specific task (e.g., translation) as in FIG. 812 in a situation where it supports only FP32 and FP16 data types. On the other hand, a second electronic device (B device) (820) may configure an AI model for performing a specific task (e.g., translation) as in FIG. 822 in a situation where it supports INT 8 and INT4 in addition to FP32. A third electronic device (C device) (830) may configure an AI model for performing a specific task (e.g., translation) as in FIG. 832 in a situation where it supports INT 8, INT4, and INT 2. Even if each electronic device is the same, the composition of the AI ​​model may vary depending on the type of task (e.g., task 1, task 2, task 3).

[0179] Each electronic device (e.g., device A, B, C) may have different configurations of the prompt execution unit (e.g., BERT model) and token generation unit (e.g., KVcache model) for the same task depending on preset guidelines or resource conditions (e.g., memory capacity). However, even if the configurations of the prompt execution unit and token generation unit are different, the inference and output results of the AI ​​model can satisfy the KPIs of each electronic device. In addition, learning and inference can be performed using a single original model on different electronic devices. Inference can refer to the process of making predictions on new data using the learned model. The AI ​​model can apply the patterns learned in the learning phase to new data and output predicted values. The results of inference can vary depending on the performance (e.g., accuracy, speed) of the AI ​​model. In Figure 834, when the third electronic device (830) processes the second task, the first data type of the prompt execution unit can be determined as INT8, and the second data type of the token generation unit can be determined as INT4.

[0180] In Figure 836, the third electronic device (830) may determine the first data type of the prompt execution unit as INT2 and the second data type of the token generation unit as INT4 when processing the third task. This is merely an example, and the first data type of the prompt execution unit and the second data type of the token generation unit may vary depending on the specifications of the third electronic device (830) and the specifications required by the task being performed.

[0181] Figure 8b illustrates an embodiment of the present invention in which the data types of the prompt execution unit and the token generation unit are determined differently in an artificial intelligence model that performs inference by stacking the same layers.

[0182] According to one embodiment, when using quantized models with large data sizes, the electronic device can achieve KPIs that are relatively close to the original model, compared to when configuring an AI model using a single data type. Furthermore, when using quantized models with large data sizes, the electronic device can save on RAM or ROM because the data size is relatively smaller.

[0183] In one embodiment, even if the configuration of the prompt execution unit and token generation unit differs, the inference and output results of the AI ​​model can satisfy the KPIs required by each electronic device. It is also possible to perform learning and inference using a single source model on different electronic devices, achieving consistent KPIs.

[0184] The original AI model (840) can perform inference by constructing a neural network by stacking multiple layers having a single data type (e.g., FP32). While the data type is assumed to be FP32 in FIG. 8b, the data types that the original AI model (840) can have are not limited to this. However, the original AI model (840) can set all of its multiple layers to have the same data type.

[0185] On the other hand, the electronic device according to this document (e.g., the electronic device (300) of FIG. 3) can determine that each layer of an artificial intelligence model (850) that constitutes a neural network by stacking multiple layers has a different data type. In addition, not only the layers having a stack structure but also each operation can be determined to have various independent data types. The electronic device (300) can test the output of the artificial intelligence model using various data types not only in the layers but also in the process of performing the operation. The electronic device (300) can test the output of the artificial intelligence model in the process of performing the operation to determine the data type with the best performance. In other words, multiple operations may also have different data types rather than a single data type.

[0186] For example, the electronic device (300) may determine to have one layer have a data type of INT8 and another layer have a data type of INT4. The electronic device (300) may determine to have another layer have a data type of FP16. Multiple layers may be combined to form a transformer. The transformer may learn sentences by using an attention mechanism to focus more on important words and less on less important words. The original artificial intelligence model (840) according to the comparative embodiment may have a form in which layers having the same kind of data type are stacked. On the other hand, the artificial intelligence model (850) according to various embodiments of the present document may have a form in which layers having different kinds of data types are stacked.

[0187] In one embodiment, inference can refer to the process of making predictions about new data using a trained model. An AI model can apply patterns learned during the training phase to new data and output predicted values. The results of inference can vary depending on the performance (e.g., accuracy, speed) of the AI ​​model.

[0188] According to one embodiment, when the type of task to be executed is determined and there are preset conditions for the first data and the second data, the electronic device (300) can determine the types of the first data and the second data based on the preset conditions, and configure the artificial intelligence model and perform inference based on the determined data types.

[0189] According to one embodiment, when the type of task to be executed is determined but there are no preset conditions for the first data and the second data, the electronic device (300) may measure the performance of each combination of data by making inferences using the artificial intelligence model while changing the types of the first data and the second data. The electronic device (300) may configure the artificial intelligence model with a combination that has relatively the best performance. The performance of each combination of data may include at least one of the size of the memory, the memory capacity (capability) of the backend, the operation speed (key performance indicators, KPI), power consumption, or inference accuracy. The KPI may refer to the number of tokens processed per second (tokens / sec) by the BERT model or the Kvcache model. For example, the more tokens the BERT model processes per second, the faster the processing speed of the artificial intelligence model.

[0190] FIG. 9 is a flowchart illustrating a method for an electronic device to configure an artificial intelligence model according to various embodiments.

[0191] The operations described through FIG. 9 may be implemented based on instructions that may be stored in a computer recording medium or memory (e.g., memory (130) of FIG. 1). The illustrated method (900) may be executed by an electronic device (e.g., electronic device (300) of FIG. 3) described above through FIGS. 1 to 8B, and the technical features described above will be omitted below. The order of each operation of FIG. 9 may be changed, some operations may be omitted, and some operations may be performed simultaneously.

[0192] In operation 910, the electronic device (300) or processor (e.g., the processor (310) of FIG. 3) may determine a type of first data to be provided to the prompt execution unit from among a plurality of data types. The prompt execution unit may refer to a configuration that executes a user prompt. The prompt execution unit may include, for example, a BERT (bidirectional encoder representations from transformers) model as a model for processing the input prompt.

[0193] In operation 920, the electronic device (300) may determine the type of second data to be provided to the token generation unit. The token generation unit may include a configuration for generating a single token. The token generation unit may include a key-value cache (KVcache). A key-value cache may refer to a cache system that stores intermediate results to perform a specific task relatively quickly. The token generation unit (524) stores a key and a corresponding value as data, and can quickly find the value using the key.

[0194] In operation 930, the electronic device (300) may generate an artificial intelligence model based on a first data type and a second data type. The existing artificial intelligence model may be referred to as a first artificial intelligence model. The artificial intelligence model generated based on the first data type and the second data type may be referred to as a second artificial intelligence model. The first data type may refer to a data type of a prompt execution unit. The second data type may refer to a data type of a token generation unit. According to one embodiment, the data types may include FP32, FP16, INT8, INT4, and INT2.

[0195] At operation 940, the electronic device (300) can evaluate the performance of the newly generated second artificial intelligence model and decide whether to use it.

[0196] The first AI model can store at least one of accuracy, memory usage, or computational speed in memory. The accuracy of the AI ​​model can be determined based on the percentage of output results that match the correct answer. The correct answer may vary depending on the data set input to the AI ​​model. For example, when performing training to distinguish between cats and dogs, each image may be labeled with "cat" or "puppy." The accuracy of the AI ​​model can be determined based on the percentage of matches between the AI ​​model's predicted values ​​and the actual labels.

[0197] The first AI model can compare at least one of the following metrics: accuracy, memory usage, or computational speed. For example, if the accuracy of the second AI model is relatively higher than that of the first AI model, the first AI model may decide to use the newly generated second AI model. Accuracy, memory usage, or computational speed are merely examples, and the metrics for comparing performance between AI models are not limited to these.

[0198] According to one embodiment, the electronic device (300) may determine the type of first data to be provided to the prompt execution unit from among a plurality of data types by a table stored in the memory, and may determine the type of second data to be provided to the token generation unit. The table may include a table in which accuracy is scored for each model when a plurality of models are composed based on the type of the first data and the second data.

[0199] According to one embodiment, the electronic device (300) may determine the type of first data to be provided to the prompt execution unit and the type of second data to be provided to the token generation unit among a plurality of data types for each task or operation based on a table in which accuracy is scored for each model. The task or operation may include any one of language translator, summarization, typo correction, or reply. The task or operation is merely an example and is not limited thereto.

[0200] According to one embodiment, the electronic device (300) can check the resources of the electronic device if there is no table corresponding to the task to be performed. The electronic device (300) can measure the memory usage and the operation speed for each model when a plurality of models are composed based on the type of the first data and the second data based on whether the remaining resources exceed a designated level. The electronic device (300) can determine a combination having the fastest operation speed among combinations of data types whose memory usage is relatively less than the remaining resources of the electronic device, and perform inference and the task using a model composed of the determined combination.

[0201] According to one embodiment, when the type of task to be executed is determined and there are preset conditions for the first data and the second data, the electronic device (300) can determine the types of the first data and the second data based on the preset conditions, and configure the artificial intelligence model and perform inference based on the determined data types.

[0202] According to one embodiment, when the type of task to be executed is determined but there are no conditions set in advance for the first data and the second data, the electronic device (300) can measure the performance of each combination of data by inferring through the artificial intelligence model while changing the types of the first data and the second data. The electronic device (300) can configure the artificial intelligence model with the combination that has the relatively best performance. The performance of each combination of data may include at least one of the size of the memory, the memory capacity (capability) of the backend, the operation speed (KPI, key performance indicators), power consumption, or the accuracy of the inference.

[0203] FIG. 10 is a flowchart illustrating a method for an electronic device to configure an artificial intelligence model according to various embodiments.

[0204] The operations described through FIG. 10 may be implemented based on instructions that may be stored in a computer recording medium or memory (e.g., memory (130) of FIG. 1). The illustrated method (1000) may be executed by an electronic device (e.g., electronic device (300) of FIG. 3) described above through FIGS. 1 to 8B, and the technical features described above will be omitted below. The order of each operation of FIG. 10 may be changed, some operations may be omitted, and some operations may be performed simultaneously.

[0205] In operation 1010, the electronic device (300) can check whether a preset manual exists for creating an artificial intelligence model. If a combination of data types that satisfies the specifications (e.g., computational speed, memory capacity, etc.) required for each task exists, the electronic device (300) can set the manual to change to that combination. Alternatively, the manual can be changed according to user settings.

[0206] In operation 1012, the electronic device (300) may configure an AI model according to a preset manual based on the existence of a preset manual for creating an AI model. The AI ​​model may include a prompt execution unit and a token generation unit. The operation of configuring the AI ​​model may refer to an operation of determining the data type used in the prompt execution unit and the token generation unit.

[0207] In operation 1014, the electronic device (300) may check the resources of the electronic device (300) based on the absence of a preset manual for creating an artificial intelligence model. The resources of the electronic device (300) may include, for example, the capacity of a CPU or the capacity of a memory.

[0208] In operation 1020, the electronic device (300) may determine whether the remaining resources exceed a specified level. The specified level may vary depending on, for example, the specifications required for each task (e.g., computational speed, memory capacity, etc.). If the specifications required for each task (or task) are high, it may be a high-level task, in which case a relatively large amount of remaining resources may be required. The electronic device (300) may set the specified level relatively high for tasks requiring high specifications, thereby maintaining the AI ​​model without changing it only when the remaining resources are relatively large. Conversely, if the remaining resources are relatively small, the AI ​​model may be changed. The specified level may vary depending on user settings.

[0209] In operation 1024, the electronic device (300) may check the remaining resources and performance while changing the data type of the artificial intelligence model based on whether the remaining resources exceed a specified level. The performance may include, for example, at least one of memory size, backend memory capacity, computational speed (key performance indicators, KPIs), power consumption, or inference accuracy.

[0210] In operation 1026, the electronic device (300) determines a data type based on remaining resources and performance, and may create a new optimal model or maintain an existing model. The operation of creating an optimal model may refer to an operation of determining a data type used in the prompt execution unit and token generation unit that is different from the original AI model.

[0211] In operation 1022, the electronic device (300) may terminate the operation without changing the data type of the artificial intelligence model based on the remaining resources being below a specified level.

[0212] According to one embodiment, an electronic device includes a memory storing a first artificial intelligence model, and a processor, wherein the first artificial intelligence model includes a prompt execution unit that executes a user prompt, and a token generation unit that generates a single token, wherein the processor measures at least one of accuracy, memory usage, and operation speed of the first artificial intelligence model based on execution of instructions stored in the memory and stores the measured results in the memory, determines a type of first data to be provided to the prompt execution unit among a plurality of data types, determines a type of second data to be provided to the token generation unit, generates a second artificial intelligence model based on the type of the first data and the type of the second data, and compares at least one of accuracy, memory usage, and operation speed of the second artificial intelligence model with the first artificial intelligence model to determine whether to use the second artificial intelligence model.

[0213] According to one embodiment, the plurality of data types may include FP32, FP16, INT8, INT4, and INT2.

[0214] According to one embodiment, the processor may determine a type of first data to be provided to the prompt execution unit from among a plurality of data types by a table stored in the memory, and may determine a type of second data to be provided to the token generation unit. The table may include a table in which accuracy is scored for each model when a plurality of models are composed based on the type of the first data and the second data, and the accuracy of the first artificial intelligence model and the second artificial intelligence model may be determined based on the ratio of results that match the correct answer among the output results.

[0215] According to one embodiment, the processor may determine a type of first data to be provided to the prompt execution unit from among a plurality of data types by a table stored in the memory, and may determine a type of second data to be provided to the token generation unit. The table may include a table in which accuracy is scored for each model when a plurality of models are composed based on the type of the first data and the second data, and the accuracy of the first artificial intelligence model and the second artificial intelligence model may be determined based on the ratio of results that match the correct answer among the output results.

[0216] In one embodiment, the processor may determine a first data type to be provided to the prompt execution unit and a second data type to be provided to the token generation unit based on a table that scores accuracy for each model, among a plurality of data types, for each task. The task may include any one of language translator, summarization, typo correction, or reply.

[0217] According to one embodiment, the processor may check the resources of the electronic device when there is no table corresponding to the task to be performed, measure the memory usage and the operation speed for each model when a plurality of models are composed based on the type of first data and the second data based on whether the remaining resources of the electronic device exceed a specified level, determine a combination having the fastest operation speed among combinations of data types having relatively less memory usage than the remaining resources of the electronic device, and perform inference and the task using the model composed of the determined combination.

[0218] According to one embodiment, when a type of task to be executed is determined and there are preset conditions for first data and second data, the processor can determine the types of the first data and second data based on the preset conditions, and configure an artificial intelligence model and perform inference based on the determined data types.

[0219] According to one embodiment, when the type of task to be executed is determined but there are no pre-set conditions for the first data and the second data, the processor may measure the performance of each combination of data by inferring through an artificial intelligence model while changing the types of the first data and the second data, and configure the artificial intelligence model with the combination with the relatively best performance. The performance of each combination of data may include at least one of the size of the memory, the memory capacity (capability) of the backend, key performance indicators (KPIs), power consumption, or inference accuracy.

[0220] The embodiments of this document disclosed in this specification and drawings are merely specific examples to easily explain the technical contents according to the embodiments of this document and to help understand the embodiments of this document, and are not intended to limit the scope of the embodiments of this document. Therefore, the scope of one embodiment of this document should be interpreted to include all changes or modified forms derived based on the technical idea of ​​one embodiment of this document, in addition to the embodiments disclosed herein.

Claims

1. In electronic devices, Memory for storing instructions and the first artificial intelligence model; and comprising at least one processor, The above first artificial intelligence model Prompt executor that executes the user prompt; and Includes a token generation unit that generates a single token, The above instructions, when executed by the at least one processor, cause the electronic device to Measure at least one of the accuracy, memory usage, or computational speed of the first artificial intelligence model and store it in the memory, Determine the type of the first data to be provided to the prompt execution unit among multiple data types, Determines the type of second data to be provided to the above token generation unit, Generate a second artificial intelligence model based on the type of the first data and the type of the second data, An electronic device that controls whether to use the second artificial intelligence model by comparing at least one of the accuracy, memory usage, or computational speed of the second artificial intelligence model with the first artificial intelligence model.

2. In paragraph 1, The above multiple data types are Electronic devices including FP32, FP16, INT8, INT4 and INT2.

3. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to store information in a table stored in the memory. Determine the type of the first data to be provided to the prompt execution unit among multiple data types, Controls to determine the type of second data to be provided to the above token generation unit, The above table When a plurality of models are composed based on the type of the first data and the second data, a table is included that scores the accuracy of each model. An electronic device in which the accuracy of the first artificial intelligence model and the second artificial intelligence model is determined based on the ratio of results that match the correct answer among the output results.

4. In paragraph 3, The above instructions, when executed by the at least one processor, cause the electronic device to determine the type of the first data to be provided to the prompt execution unit among a plurality of data types for each task based on a table in which accuracy is scored for each model, and Controls to determine the type of second data to be provided to the above token generation unit, The above task is An electronic device comprising any one of a language translator, summarization, typo correction, or reply.

5. In paragraph 3, The above instructions, when executed by the at least one processor, cause the electronic device to If there is no table corresponding to the task to be performed, check the resources of the electronic device, When a plurality of models are composed based on the type of the first data and the second data based on the remaining resources of the electronic device exceeding a specified level, the memory usage and the operation speed are measured for each model, Among the combinations of data types whose memory usage is relatively less than the remaining resources of the electronic device, the combination with the fastest computational speed is determined, An electronic device that controls the performance of inference and tasks using models composed of determined combinations.

6. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to When the type of task to be executed is determined and there are conditions set in advance for the first data and the second data, the types of the first data and the second data are determined based on the conditions set in advance, An electronic device that controls the configuration of the artificial intelligence model and inference according to a determined data type.

7. In paragraph 1, The above instructions, when executed by the at least one processor, cause the electronic device to If the type of task to be executed has been determined, but there are no conditions set in advance for the first data and the second data, By changing the types of the first data and the second data, the performance of each combination of data is measured by inferring through the artificial intelligence model. Controls the configuration of the above artificial intelligence model with the combination that has the best performance relative to others, Performance by combination of data An electronic device comprising at least one of the size of the memory, the capability of the backend, key performance indicators (KPI), power consumption, or inference accuracy.

8. In terms of operation method, An action of measuring at least one of the accuracy, memory usage, or computational speed of a first artificial intelligence model and storing it in memory; An action for determining the type of first data to be provided to the prompt executor among multiple data types; An action that determines the type of second data to be provided to the token generation unit; An operation of generating a second artificial intelligence model based on the type of the first data and the type of the second data; and A method comprising an action of comparing at least one of accuracy, memory usage, or computational speed of the generated second artificial intelligence model with the first artificial intelligence model to determine whether to use the second artificial intelligence model.

9. In paragraph 8, The above multiple data types are Methods including FP32, FP16, INT8, INT4 and INT2.

10. In paragraph 8, An operation of determining the type of the first data to be provided to the prompt execution unit from among a plurality of data types by a table stored in memory; Further comprising an action for determining the type of second data to be provided to the token generation unit, The above table When a plurality of models are composed based on the type of the first data and the second data, a table is included that scores the accuracy of each model. A method in which the accuracy of the first artificial intelligence model and the second artificial intelligence model is determined based on the ratio of results that match the correct answer among the output results.

11. In paragraph 10, An operation of determining the type of the first data to be provided to the prompt execution unit among multiple data types for each task based on a table in which accuracy is scored for each model; and Further comprising an action for determining the type of second data to be provided to the token generation unit, The above task is A method that includes any of language translator, summarization, typo correction, or reply.

12. In paragraph 10, An action to check the resources of an electronic device when there is no table corresponding to the task to be performed; An operation of measuring the memory usage and the operation speed for each model when a plurality of models are composed based on the type of the first data and the second data based on the remaining resources of the electronic device exceeding a specified level; An operation for determining a combination of data types having the fastest operation speed among combinations of data types whose memory usage is relatively less than the remaining resources of the electronic device; and A method further comprising performing inference and tasks using a model composed of determined combinations.

13. In paragraph 8, An operation of determining the type of the first data and the second data based on the pre-set condition when the type of the task to be executed is determined and there is a pre-set condition for the first data and the second data; and A method further comprising the action of configuring the artificial intelligence model and performing inference according to the determined data type.

14. In paragraph 8, When the type of task to be executed is determined but there is no condition set in advance for the first data and the second data, an operation of measuring performance for each combination of data by performing inference through an artificial intelligence model while changing the type of the first data and the second data; and It further includes the operation of reconfiguring the above artificial intelligence model with the relatively best performing combination, Performance by combination of data A method comprising at least one of memory size, backend capability, key performance indicators (KPI), power consumption, or inference accuracy.

15. In recording media, Memory for storing the first artificial intelligence model; and Contains a processor, The above first artificial intelligence model Prompt executor that executes the user prompt; and Includes a token generation unit that generates a single token, The above memory, when executed, is used by the processor Measure at least one of the accuracy, memory usage, or computational speed of the first artificial intelligence model and store it in the memory, Determine the type of the first data to be provided to the prompt execution unit among multiple data types, Determines the type of second data to be provided to the above token generation unit, Generate a second artificial intelligence model based on the type of the first data and the type of the second data, A recording medium storing instructions for controlling whether to use the second artificial intelligence model by comparing at least one of the accuracy, memory usage, or computational speed of the second artificial intelligence model with the first artificial intelligence model.

Citation Information

Patent Citations

  • Intelligent task discovery

    KR1020170140079A

  • Method and system for blind connection

    KR1020240024756A

  • Apparatus and method for detecting wiretapping device using automatic voice detection and frequency

    KR102437054B1

  • Device and method for providing benchmark result of artificial intelligence based model

    KR102587263B1

  • Apartment firefighting and smoke removing apparatus

    KR102702788B1