Electronic device, method, and non-transitory computer-readable storage medium for providing response to voice signal

WO2026205735A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/001532
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-28
Filing Date
2026-01-26
Publication Date
2026-10-01

Smart Images

  • Figure KR2026001532_01102026_PF_FP_ABST
    Figure KR2026001532_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise a microphone, a memory for storing instructions, and at least one processor including processing circuitry. The instructions, when executed individually or collectively by the at least one processor, may instruct the electronic device to: receive an audio input via the microphone; obtain information about the tone of a voice signal included in the audio input and user information collected within a time period defined according to the reception of the audio input; identify emotion information on the basis of the information about the tone of the voice signal and the user information; determine, on the basis of the emotion information, at least one target artificial intelligence model from among a plurality of candidate artificial intelligence models; and provide a response to the voice signal on the basis of the provision of the information about the voice signal to the at least one target artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable storage medium for providing a response to a voice signal

[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for providing a response to a voice signal.

[0002] With the advancement of electronic devices, technological development related to electronic devices equipped with artificial intelligence (AI) technology has recently been underway. Electronic devices equipped with AI technology can provide various services to users.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] An electronic device is provided. The electronic device may include a microphone. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include at least one processor comprising processing circuitry. The instructions may cause the electronic device to receive an audio input through the microphone when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain information regarding the tone of a voice signal included in the audio input and user information collected within a time interval defined by the reception of the audio input when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to identify sentiment information or emotional information based on the information regarding the tone of the voice signal and the user information when executed individually or collectively by the at least one processor. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to determine at least one target artificial intelligence model among a plurality of candidate artificial intelligence models based on the emotion information. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to provide a response to the voice signal based on providing information about the voice signal to the at least one target artificial intelligence model.

[0005] A method is provided. The method may be performed in an electronic device having a microphone. The method may include the operation of receiving an audio input through the microphone. The method may include the operation of obtaining information regarding the tone of a voice signal included in the audio input and user information collected within a time interval defined according to the reception of the audio input. The method may include the operation of identifying emotion information based on the information regarding the tone of the voice signal and the user information. The method may include the operation of determining at least one target artificial intelligence model among a plurality of candidate artificial intelligence models based on the emotion information. The method may include the operation of providing a response to the voice signal based on providing information regarding the voice signal to the at least one target artificial intelligence model.

[0006] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to receive an audio input through the microphone when executed by the electronic device having a microphone. The one or more programs may include instructions that cause the electronic device to obtain information about the tone of a voice signal included in the audio input and user information collected within a time interval defined by the reception of the audio input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify emotion information based on the information about the tone of the voice signal and the user information when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine at least one target artificial intelligence model among a plurality of candidate artificial intelligence models based on the emotion information when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to provide a response to the voice signal based on providing information about the voice signal to the at least one target artificial intelligence model when executed by the electronic device.

[0007] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.

[0008] Figure 2 shows a block diagram of a simplified electronic device.

[0009] FIG. 3 illustrates an example of a component of an electronic device for providing a response to a voice signal included in an audio input.

[0010] FIGS. 4a and 4b illustrate examples of electronic devices that determine a target model among a plurality of candidate models based on sentiment information or emotional information.

[0011] Figure 5 illustrates an example of an electronic device that changes a target model based on a change in emotional information.

[0012] FIGS. 6a and 6b illustrate examples of operations of an electronic device that provides a response to a voice signal.

[0013] Figure 7 illustrates an example of an electronic device that adaptively outputs a response to a voice signal according to emotional information.

[0014] FIGS. 8a and 8b illustrate examples of electronic devices that adaptively control an external electronic device according to emotional information.

[0015] FIGS. 9a and 9b illustrate examples of electronic devices that adaptively determine content based on emotional information.

[0016] Figure 10 is a schematic diagram of an exemplary artificial intelligence (AI) system.

[0017] Throughout the drawings, the same reference numerals will be understood to refer to the same parts, components, and structures.

[0018] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of other embodiments. A singular expression may include a plural expression unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art described in this disclosure. Terms used in this disclosure that are defined in a general dictionary may be interpreted as having the same or similar meaning as they have in the context of the relevant technology, and are not to be interpreted in an ideal or overly formal sense unless explicitly defined in this disclosure. In some cases, even terms defined in this disclosure are not to be interpreted to exclude the embodiments of this disclosure.

[0019] In the various embodiments of the present disclosure described below, a hardware-based approach is described as an example. However, since the various embodiments of the present disclosure include techniques using both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0020] Terms used in the following description to refer to data (e.g., data, information, emotion information, emotion estimation information, biometric information, signal, task), terms referring to values ​​(e.g., threshold, reference range), terms for operation states (e.g., operation, process), terms referring to objects (e.g., visual object, UI (user interface) object), terms referring to network entities, terms referring to device components, etc., are provided as examples for the convenience of explanation. Accordingly, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used. Furthermore, terms such as '...part', '...device', '...object', '...body' used below may refer to at least one shape structure or a unit that processes a function.

[0021] Additionally, in this disclosure, expressions of "greater than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled; however, this is merely for the purpose of expressing an example and does not exclude descriptions of "greater than" or "less than." Conditions described as "greater than" may be replaced with "greater than," conditions described as "less than" may be replaced with "less than," and conditions described as "greater than and less than" may be replaced with "greater than and less than." Furthermore, "A" to "B" below refer to at least one of elements from A (including A) to B (including B). Below, "C" and / or "D" refers to including at least one of "C" or "D," i.e., {"C", "D", "C" and "D"}.

[0022] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.

[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0024] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0025] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0026] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0027] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0028] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0029] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0030] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0031] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0032] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0033] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0034] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0035] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0036] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0038] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0039] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0041] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0043] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0045] In this disclosure, technology related to artificial intelligence (or artificial intelligence models) may be described. In the embodiments of this disclosure, an electronic device (e.g., electronic device (101)) may utilize an artificial intelligence model. The artificial intelligence model may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values ​​and performs neural network operations through operations between the results of operations of a previous layer and the plurality of weights. The plurality of weights possessed by the plurality of neural network layers may be optimized by the learning results of the artificial intelligence model. For example, the plurality of weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. The artificial neural network may include a Deep Neural Network (DNN). For example, artificial neural networks may include, but are not limited to, CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), BRDNN (Bidirectional Recurrent Deep Neural Network) and / or Deep Q-Networks.

[0046] The electronic device (101) may provide a virtual assistant service using artificial intelligence (or an artificial intelligence model). For example, providing a virtual assistant service may include providing the functions of a virtual assistant. In the present disclosure, a virtual assistant may be referred to by technical terms such as voice assistant, digital assistant, intelligent automated assistant, automatic digital assistant, artificial intelligence assistant, intelligent assistant, personal assistant, mobile assistant, intelligent agent, and / or equivalent. By example, without limitation, a virtual assistant may include Bixby. However, the present disclosure is not limited thereto.

[0047] A virtual assistant may be a software application (or software agent) that processes a task requested by a user (or user input) of an electronic device (101) and provides a service. A task may be referred to as an operation (or action) that the electronic device (101) must perform using the virtual assistant. For example, the electronic device (101) may provide a virtual assistant service by executing the software application. The electronic device (101) providing the virtual assistant service may identify or determine a task (or function) requested by the user based on user input. The electronic device (101) may provide a response based on user input by executing the identified task. For example, user input may include a voice signal. For example, the electronic device (101) may obtain text information regarding the voice signal by performing natural language processing on the voice signal.

[0048] According to one embodiment, an electronic device (101) providing a virtual assistant service can provide a response to a voice signal by executing a pre-designated task corresponding to text information. For example, the electronic device (101) may have a list of pre-designated tasks corresponding to text information. For example, the electronic device (101) may use the list of pre-designated tasks to identify a matching level between each of the pre-designated tasks and the text information. For example, the matching level may be referenced as a corresponding degree and / or similarity. For example, the electronic device (101) may use the list of pre-designated tasks to identify or execute the task with the highest matching level among the pre-designated tasks. For example, even if the electronic device (101) obtains a user input that is the same (or similar) to a previous user input, it may provide a response identical to a previous response because it identifies the task independently of the previously identified task. For example, the first emotion information corresponding to the previous user input and the second emotion information corresponding to the user input may be different. For example, even though the emotional state indicated by the first emotional information and the emotional state indicated by the second emotional information are different, the electronic device (101) may provide the same response as the previous response. For example, the user experience of the electronic device (101) that provides a response to a voice signal independently of the emotional information (or emotional state) may be lower than the user experience of the electronic device (101) that provides a response to a voice signal adaptively according to the emotional state. Accordingly, a method of providing a virtual assistant service that provides a response adaptively according to the user's emotional information may be required. For example, the quality of a virtual assistant service that provides a response adaptively according to the user's emotional information may be higher than the quality of a virtual assistant service that provides a response independently of the user's emotional information.

[0049] The present disclosure may describe a virtual assistant service that adaptively provides a response based on the user's emotional information. In embodiments of the present disclosure, an electronic device (101) may determine a target model among a plurality of candidate models based on emotional information. For example, the plurality of candidate models may include a plurality of artificial intelligence models. In the present disclosure, a model may be understood and referred to as an artificial intelligence model. For example, the plurality of candidate models may be referred to as a plurality of artificial intelligence models available to the electronic device (101). For example, some of the plurality of candidate models may be included in the electronic device (101), and other parts of the plurality of candidate models may be included in an external electronic device (e.g., electronic device (102), electronic device (104), server (108)) distinct from the electronic device (101). For example, the target model may be referred to as an artificial intelligence model determined to perform inference on a voice signal obtained by the electronic device (101). In the present disclosure, the target model may be referred to as a target artificial intelligence model. For example, the electronic device (101) may provide a response to a voice signal by providing information about the voice signal (e.g., a prompt indicating the voice signal) to the target model. For example, the electronic device (101) may provide a response to the voice signal in a manner according to emotional information. For example, components for providing such a method may be included in the electronic device (101). Such components will be described and illustrated with reference to FIGS. 2 and 3.

[0050] FIG. 2 illustrates a block diagram of a simplified electronic device (101). The electronic device (101) may be one of various types of mobile devices, such as a laptop, smartphones with various form factors (e.g., bar-type smartphones, foldable-type smartphones, multi-foldable-type smartphones, or rollable-type smartphones), a tablet, a cellular phone, and other similar computing devices. The electronic device (101) may be referred to as a user device, a multi-functional device, or a portable device.

[0051] Referring to FIG. 2, the electronic device (101) may include at least one processor (200), memory (210), display (220), communication circuit (230), speaker (240), and / or microphone (250). For example, at least one processor (200), memory (210), display (220), communication circuit (230), speaker (240), and / or microphone (250) may be electrically and / or operably coupled with each other by a communication bus.

[0052] In the following, the hardware components being operatively coupled may mean that a direct or indirect connection between the hardware components is established via wired or wireless means so that a second hardware component is controlled by a first hardware component among the hardware components. Although the hardware components illustrated in FIG. 2 are illustrated based on different blocks, the present disclosure is not limited thereto. For example, some of the hardware components illustrated in FIG. 2 (e.g., at least one processor (200), memory (210), and at least a portion of a communication circuit (230)) may be included in a single integrated circuit such as a system on chip (SoC) or a system in package (SIP). The type and / or number of hardware components included in the electronic device (101) are not limited to those illustrated in FIG. 2. For example, the electronic device (101) may include only some of the hardware components illustrated in FIG. 2.

[0053] At least one processor (200) may include a hardware component for processing data based on executing instructions. At least one processor (200) may be configured to execute instructions stored in memory (210) individually or collectively. At least one processor (200) may include a processing circuit. For example, the hardware component for processing data may include an arithmetic and logic unit (ALU), a floating point unit (FPU), and a field programmable gate array (FPGA). For example, the hardware component for processing data may include a central processing unit (CPU), a graphic processing unit (GPU), a display processing unit (DPU), a neural processing unit (NPU), a digital signal processor (DSP), an application processor (AP), and / or a microcontroller (MCU). At least one processor (200) may include one or more cores. For example, at least one processor (200) may have the structure of a multi-core processor such as a dual core, quad core, or hexa core. The at least one processor (200) of FIG. 2 may have substantially the same content as the processor (120) of FIG. 1.

[0054] Memory (210) may include a hardware component for storing data and / or instructions that are input to and / or output from at least one processor (200). For example, instructions may represent operations and / or actions to be performed on the data by at least one processor (200) of the electronic device (101). Instructions may be referred to as programs, firmware, operating systems, processes, routines, sub-routines, and / or applications. Memory (210) may include one or more storage media. Memory (210) may include volatile memory, such as random-access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM). Volatile memory may include at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, or pseudo SRAM (PSRAM). Non-volatile memory may include, for example, at least one of PROM (programmable ROM), EPROM (erasable PROM), EEPROM (electrically erasable PROM), flash memory, hard disk, compact disk, or EMMC (embedded multimedia card). The specific details regarding the memory (210) of FIG. 2 may be substantially the same as the details regarding the memory (130) of FIG. 1.

[0055] The display (220) may include hardware components of an electronic device (101) used to display a screen. For example, the display (220) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light-emitting diode (OLED) or a micro LED. However, it is not limited thereto. For example, the display (220) may include a liquid crystal display (LCD).

[0056] The communication circuit (230) may include hardware components to support the transmission and / or reception of signals between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), server (108)). The communication circuit (310) may include, for example, at least one of a modem, an antenna, or an O / E (optic / electronic) converter. The communication circuit (310) may support the transmission and / or reception of electrical signals based on various types of protocols such as Ethernet, LAN (local area network), WAN (wide area network), WiFi (wireless fidelity), Bluetooth, BLE (Bluetooth low energy), Zigbee, LTE (long term evolution), and 5G NR (new radio). Specific details regarding the communication circuit (230) of FIG. 2 may be substantially the same as the communication module (190) and / or antenna module (197) of FIG. 1.

[0057] A speaker (240) may be used to output (or provide) audio (or sound) configured through at least one processor (200). For example, an electronic device (101) may output audio to the outside of the electronic device (101) through the speaker (240). Specific details regarding the speaker (240) of FIG. 2 may be substantially the same as those regarding the sound output module (155) of FIG. 1.

[0058] The microphone (250) may include a hardware component of an electronic device (101) used to acquire audio input. For example, the microphone (250) may be configured to acquire audio input generated around the electronic device (101). For example, the audio input may include a user's voice signal. The specific details regarding the microphone (250) of FIG. 2 may be substantially the same as the details regarding the input module (150) of FIG. 1.

[0059] FIG. 3 illustrates an example of a component of an electronic device for providing a response to a voice signal included in an audio input.

[0060] Referring to FIG. 3, the electronic device (101) may include an audio front end (301), an automatic speech recognition (ASR) module (303), a tone identification module (305), an environment identification module (307), a user information collection module (309), an emotion estimation information identification module (311), a sensing information management module (315), an emotion information identification module (317), a target model determination module (319), model information (321), a task execution module (327), a response generation module (329), and / or a response provision module (331). However, the embodiments are not limited to the modules illustrated in FIG. 3. The modules of the electronic device (101) are not limited to the embodiments, and at least one module may be integrated or additional modules may be included to perform operations, and at least one module may be implemented in hardware or software.

[0061] The audio front end (301) may be used to perform preprocessing on an audio input obtained through a microphone (e.g., microphone (250)) so that the electronic device (101) modules can use it. For example, the preprocessing performed on the audio input may include noise removal (or noise canceling), normalization, amplification, and / or filtering.

[0062] The audio front end (301) can be used to perform voice segmentation on the audio input. For example, the electronic device (101) can identify voice segments and non-voice segments of the audio input by performing voice segmentation on the audio input. For example, the electronic device (101) can identify segments corresponding to the voice (or speech) of a person (e.g., user of the electronic device (101)) by performing voice segmentation (e.g., voice activity detection) on the audio input.

[0063] The audio front end (301) may be used to identify or extract features from the audio input. The audio front end (301) may be used to perform processing to improve the quality of the audio input. For example, processing to improve the quality of the audio input may include echo removal and / or noise suppression.

[0064] The audio front end (301) may be used to identify a voice signal included in the audio input. For example, the voice signal may represent the voice (or speech) of a user of a person (e.g., the electronic device (101)). For example, the voice signal may be referred to as a voice signal, a sound signal, a signal, a speech signal, and / or a user signal. For example, the audio front end (301) may be used to identify a background noise signal included in the audio input. For example, the background noise signal may include a background sound. For example, the electronic device (101) may use the audio front end (301) to separate the audio input into a voice signal and a background noise signal. For example, separating the audio input into a voice signal and a background noise signal may be referred to as audio source separation.

[0065] The audio front end (301) may provide a portion of the audio input to the ASR module (303), the tone identification module (305), and / or the environment identification module (307). The audio front end (301) may perform processing on the audio input to provide a portion of the audio input to each of the ASR module (303), the tone identification module (305), and / or the environment identification module (307). For example, processing on the audio input may be performed in different ways depending on the module to which the audio input is provided. For example, the audio front end (301) may perform noise canceling (or noise removal) on the audio input. For example, the audio front end (301) may provide the voice signal contained in the audio input to the ASR module (303) by performing noise canceling on the audio input. For example, the audio front end (301) may provide the voice signal contained in the audio input to the tone identification module (305) by performing noise canceling on the audio input. For example, the degree of noise canceling performed on the voice signal provided to the ASR module (303) and the degree of noise canceling performed on the voice signal provided to the tone identification module (305) may differ. For example, the audio front end (301) may provide or transmit a background noise signal included in the audio input to the environment identification module (307). For example, the audio front end (301) may provide the background noise signal to the environment identification module (307) by performing processing to extract the background noise signal from the audio input and / or amplify the background noise signal from the audio input. The audio front end (301) may request or command the user information collection module (309) to collect user information. For example, the audio front end (301) may provide a command signal related to the collection of user information to the user information collection module (309).For example, the electronic device (101) can provide a command signal related to the collection of user information to the user information collection module (309) using the audio front end (301) in response to acquiring an audio input through a microphone (e.g., microphone (250)).

[0066] As an example, but not limited to, an audio front end (301) may be used to identify a wake-up word included in an audio input. For example, an electronic device (101) may execute a function associated with the wake-up word based on identifying the wake-up word using the audio front end (301). For example, an electronic device (101) may activate a virtual assistant service based on identifying the wake-up word. For example, the wake-up word may be referred to as a wake-up voice input, a lexical trigger, a hot-phrase, a keyword, a wake word, a hot-word, a trigger word, a trigger phrase, a trigger expression, and / or an equivalent technical term.

[0067] An automatic speech recognition (ASR) module (303) can be used to convert a voice signal included in an audio input into text information. For example, the ASR module (303) can be used to identify or extract acoustic features of a voice signal. For example, an electronic device (101) can convert a voice signal into text information according to the acoustic features of the voice signal by using the ASR module (303). For example, the ASR module (303) can provide a speech-to-text (STT) function. For example, the electronic device (101) can apply STT to a voice signal by using the ASR module (303). For example, the electronic device (101) can apply STT to a voice signal based on the context of the voice signal. By applying STT to a voice signal, the electronic device (101) can obtain text information corresponding to the voice signal. For example, text information may include text representing a user's utterance expressed by a voice signal. For example, the ASR module (303) may include an artificial intelligence model for providing STT functionality. For example, the artificial intelligence model for providing STT functionality may include a language model. The ASR module (303) may provide text information to the emotion information identification module (317). The ASR module (303) may provide text information to the test execution module (327).

[0068] A tone identification module (305) may be used to identify or obtain information regarding the tone of a voice signal included in an audio input. For example, information regarding the tone of a voice signal may represent acoustic characteristics of the voice signal. For example, acoustic characteristics of the voice signal may include the loudness of the voice signal, the pitch of the voice signal, spectral information of the voice signal, the speed of the voice signal, the timbre of the voice signal, and / or the intensity of the voice signal. For example, acoustic characteristics of the voice signal may be represented by acoustic parameters including the loudness of the voice signal, the pitch of the voice signal, spectral information of the voice signal, the speed of the voice signal, the timbre of the voice signal, and / or the intensity of the voice signal. As an example, but not limited to, acoustic characteristics of the voice signal may include information obtained, derived, or generated according to the acoustic parameters. For example, information obtained according to the acoustic parameters may include information indicating the degree to which each of the acoustic parameters has changed over time, information indicating the average of each of the acoustic parameters in each time interval (e.g., each audio frame), and information indicating the variance of each of the acoustic parameters in each time interval. As an example, but not limited to, information obtained according to the acoustic parameters may include Mel-Frequency Cepstral Coefficients (MFCC) information of the voice signal and / or Cepstrum information of the voice signal.

[0069] The tone identification module (305) may be used to determine or identify emotions based on acoustic features. For example, the tone identification module (305) may include an artificial intelligence model trained to identify or determine emotions based on acoustic features by providing acoustic features. For example, the artificial intelligence model may be trained using emotions (or emotion information) associated with acoustic feature data. For example, the tone identification module (305) may output a type of emotion based on acoustic features and / or a probability for said type. For example, emotions based on acoustic features may be used to obtain emotion information obtained using the emotion information identification module (317) described later. For example, emotions based on acoustic features may be referred to as emotions represented by voice signals. The tone identification module (305) may provide information regarding the tone of the voice signal to the emotion information identification module (317). The tone identification module (305) may be referred to as a voice sentiment recognition or voice emotion recognition module.

[0070] The environment identification module (307) may be used to identify information about the environment in which the electronic device (101) is located, based on a background noise signal included in the audio input. For example, the environment may be the environment in which the electronic device (101) is located while the audio input is acquired through a microphone (e.g., microphone (250)). For example, information about the environment may be used to acquire emotional information, which will be described later. For example, the electronic device (101) may use information about the environment to identify or determine the effect that the environment in which the electronic device (101) is located has on the user's emotions. For example, information about the environment may include acoustic characteristics of the background noise signal included in the audio input.

[0071] An environment identification module (307) may be used to identify acoustic features of a background noise signal included in an audio input. For example, the environment identification module (307) may receive a background noise signal from an audio front end (301). For example, the background noise signal may include background sounds. For example, the acoustic features of the background noise signal may include the type of background sound, the distance between the source of the background sound and the electronic device (101), the intensity of the background sound, and / or the temporal variation of the background sound. For example, the type of background sound may include conversation, music, and / or natural sounds, but the embodiments are not limited. For example, the electronic device (101) may use the environment identification module (307) to identify information about the environment in which the electronic device (101) is located according to the acoustic features of the background noise signal. The environment identification module (307) may provide information about the environment in which the electronic device (101) is located to the emotion information identification module (317). For example, the environment identification module (307) can be referred to as an audio scene analysis module.

[0072] The user information collection module (309) may be used to obtain user information. For example, user information may include text identified or determined by the electronic device (101) according to user input, messages within the electronic device (101) (e.g., SMS (short message service) messages, SNS (social network service) messages, conversation data related to a message application, messages from a chat application), search history within the electronic device (101), events for notifications (e.g., advertisement notifications, schedule notifications), movement data of the electronic device (101), and / or pattern information of user input, but the embodiments are not limited. For example, movement data of the electronic device (101) may be obtained through a sensor (not shown) of the electronic device (101). For example, the sensor may include a motion sensor, an accelerometer, a geomagnetic sensor, an angular velocity sensor, and / or an IMU (inertial measurement unit) sensor. For example, movement data of the electronic device (101) may indicate shaking of the electronic device (101). For example, the shaking of the electronic device (101) may be based on the action of the user of the electronic device (101). For example, motion data of the electronic device (101) may indicate the degree of shaking, the pattern of shaking, and / or the speed of shaking, but the embodiments are not limited. For example, pattern information of user input may indicate the action of the user regarding the user input. For example, pattern information of user input may include data indicating an action related to the frequency of user input. For example, user input may include multimodal input, touch input, gesture input, scroll input, and / or swipe input, but the embodiments are not limited.

[0073] As an example not limited to, user information may include previous emotional information identified before the emotional information is identified by the electronic device (101). For example, user information may include changes in emotional information identified by the electronic device (101).

[0074] According to one embodiment, the user information collection module (309) may be used to collect user information within a defined time interval based on receiving audio input by a microphone (e.g., microphone (250)). For example, the electronic device (101) may collect user information in the background. For example, the electronic device (101) may use the user information collection module (309) to identify or obtain user information collected within a specified time interval (e.g., 30 minutes) prior to the point in time when the audio input is received. For example, user information may include voice signals (e.g., voice commands) collected within a specified time interval prior to the point in time when the audio input is received and / or conversation history with a virtual assistant collected within a specified time interval prior to the point in time when the audio input is received.

[0075] The user information collection module (309) can provide user information to the emotion estimation information identification module (311). The user information collection module (309) can be referred to as a data watcher.

[0076] The emotion estimation information identification module (311) can be used to obtain emotion estimation information based on user information. For example, the emotion estimation information identification module (311) can identify words related to emotions and / or contexts related to emotions by identifying text within user information. For example, the emotion estimation information identification module (311) can output emotion estimation information based on words related to emotions, contexts related to emotions, and / or pattern information of user input. For example, if a text message within the electronic device (101) has relatively positive content (e.g., winning an entry, passing an exam), the emotion estimation information may indicate joy. For example, if the pattern information of user input indicates a frequency of relatively short touch inputs and / or repetitive scroll inputs, the emotion estimation information may indicate tension (or stress). The emotion estimation information may be used to obtain emotion information to be described later. The emotion estimation information identification module (311) may provide emotion estimation information to the emotion information identification module (317). The emotion estimation information identification module (311) can be referred to as a user behavioral sentiment estimation or user behavioral emotion estimation module.

[0077] A sensing information management module (315) may be used to acquire and manage sensing information from a wearable device (313). For example, the wearable device (313) may include a finger-worn electronic device (e.g., a smart ring), a wrist-worn electronic device (e.g., a smart watch), and / or a head-worn electronic device (e.g., smart glasses), but the embodiments are not limited. For example, the sensing information may be acquired by at least one sensor within the wearable device (313). For example, at least one sensor within the wearable device (313) may include a heart rate sensor, an accelerometer, a gyroscope, a temperature sensor, an oxygen saturation sensor, a galvanic skin response sensor (GSR) sensor, and / or a blood pressure sensor, but the embodiments are not limited. The sensing information may include biometric information of a user of the wearable device (313) (e.g., a user of the electronic device (101)). For example, the sensing information may include heart rate, heart rate variability, blood pressure, oxygen saturation, sleep quality, body temperature, and / or step count, but the embodiments are not limited. The sensing information management module (315) may provide the sensing information to the emotion information identification module (317). For example, the sensing information management module (315) may be referred to as a sensor monitor.

[0078] The emotion information identification module (317) may be used to obtain emotion information. For example, the emotion information identification module (317) may obtain, identify, or determine emotion information based on text information corresponding to a voice signal, information about the tone of the voice signal, information about the environment in which the electronic device (101) is located, and / or sensing information. For example, the emotion information may indicate the emotional state of the speaker of the voice signal (e.g., user of the electronic device (101)) and / or a change in the speaker's emotion. For example, the emotion information may indicate a comprehensive emotion obtained based on text information corresponding to the voice signal, information about the tone of the voice signal, information about the environment in which the electronic device (101) is located, and / or sensing information. For example, the emotion information may include the cause of the emotion and / or the context of the emotion. The type of the emotion state may be referred to as the emotion type. For example, types of emotional states may include joy, sadness, anger, surprise, fear, confusion, boredom, relief, calmness, loneliness, hope, jealousy, and / or disappointment. The types of emotional states are examples for illustrative purposes only and are not limiting. Additionally, the exemplified types of emotional states may be understood as different types of emotional states depending on the intensity or degree of the corresponding type. For example, it may be assumed that the first type of emotional state and the second type of emotional state each represent joy. However, the degree of joy represented by the first type of emotional state may differ from the degree of joy represented by the second type of emotional state. If the degree of joy represented by the first type of emotional state differs from the degree of joy represented by the second type of emotional state, the first type of emotional state and the second type of emotional state may be understood as different types.

[0079] As a non-limited example, the type of emotional state may be related to psychological state. For instance, the type of emotional state may be expressed using anxiety levels, concentration levels, fatigue levels, and / or stress levels.

[0080] As a non-limiting example, the emotion information identification module (317) may be used to determine the value of a hyperparameter corresponding to the emotion information. For example, the hyperparameter may be referenced as a parameter set in the artificial intelligence model to control the inference behavior of the artificial intelligence model. The hyperparameter will be described and exemplified through the description of the target model determination module (319) to be described later.

[0081] The emotion information identification module (317) may include an artificial intelligence model. As an example, but not limited to, the emotion information identification module (317) may be a module that performs actions according to rules. The emotion information identification module (317) may provide emotion information to the target model determination module (319). The emotion information identification module (317) may be referred to as an emotion module (or emotion model).

[0082] The target model determination module (319) may be used to determine a target model among a plurality of candidate models based on emotional information. For example, the plurality of candidate models may be referred to as a plurality of artificial intelligence models. For example, the plurality of candidate models may be referred to as a plurality of artificial intelligence models available to the electronic device (101). For example, the target model may be referred to as a target artificial intelligence model. For example, the target model may be referred to as an artificial intelligence model determined to perform inference on a voice signal acquired by the electronic device (101). As an example not limited to, the target model determination module (319) may be used to determine a target model among a plurality of candidate models based further on the state of the electronic device (101) and / or the voice signal.

[0083] The target model determination module (319) may be used to change the hyperparameters of the target model determined among a plurality of candidate models. For example, the target model determination module (319) may change the hyperparameters of the target model using the value of the hyperparameter determined by the emotion information identification module (317). The hyperparameters may be referenced as parameters set in the artificial intelligence model to control the inference behavior of the artificial intelligence model. For example, the hyperparameters may include a parameter that limits the number of candidates to be considered when determining a token in the artificial intelligence model (e.g., top-k), a parameter that sets a threshold value that serves as a standard for cumulative probability when determining a token (e.g., top-p), and / or a parameter that controls the randomness of the output by adjusting the probability distribution of token selection (e.g., temperature), but the embodiments are not limited. As a non-limiting example, the electronic device (101) may use the emotion information identification module (317) to determine the value of a hyperparameter that enhances the creativity of the target model to be described later, according to emotion information corresponding to 'boredom'.

[0084] The target model determination module (319) can provide a signal (or data) representing the target model to the task execution module (327). The target model determination module (319) may be referred to as a model decision module.

[0085] The target model determination module (319) may use model information (321). Model information (321) may include information (325) of first candidate models (323) and / or second candidate models. In the present disclosure, each of the candidate models (e.g., first candidate models (323), second candidate models) may be referred to as a candidate artificial intelligence model in that it includes an artificial intelligence model. For example, each of the candidate models may be understood to include, represent, or refer to a generative artificial intelligence model (e.g., language model, large language model (LLM), large vision model (LVM), multimodal model, large multimodal model). Additionally, in the present disclosure, each of the candidate models may be used to provide a service (e.g., chatbot, virtual assistant) based on a generative artificial intelligence model. In the present disclosure, providing data (e.g., information about a voice signal, a prompt indicating a voice signal) to a model within the candidate models may include providing data to a generative artificial intelligence model and / or providing data to a service based on the generative artificial intelligence model.

[0086] The first candidate models (323) may be referred to as artificial intelligence models included in or stored in the electronic device (101). For example, each model within the first candidate models (323) may be referred to as an on-device model. The information (325) of the second candidate models may represent data regarding the second candidate models included in an external electronic device (e.g., electronic device (102), electronic device (104), server (108)) distinct from the electronic device (101). The information (325) of the second candidate models may include the size, learning method, use, reasoning ability, and / or learning purpose of each of the second candidate models within the external electronic device. For example, each model within the second candidate models may be referred to as a cloud-based model. For example, the size of the model within the first candidate models (323) may be smaller than the size of the model within the second candidate models. However, the embodiments are not limited thereto. For example, the inference quality of the models in the first candidate models (323) may be lower than the inference quality of the models in the second candidate models. However, the embodiments are not limited thereto.

[0087] Multiple candidate models may include a starter model, an affordable model, a midrange model, and / or a premium model. A starter model may be represented as a model that provides basic functionality but has a relatively small size. For example, a starter model may be used for performing relatively simple question answering and basic tasks. For example, a starter model may include a language model having about 1 billion to 2 billion parameters.

[0088] A cost-effective model may be a model that has a higher inference quality than the starter model. For example, a cost-effective model may have a better understanding of various situations than the starter model. For example, a cost-effective model may support multimodal capabilities. For example, a cost-effective model may include a multimodal model with approximately 3B parameters.

[0089] An intermediate model may be a model that has a higher inference quality than that of a cost-effective model. For example, because an intermediate model has an understanding of relatively diverse domains, its contextual interpretation ability may be higher than that of a cost-effective model. For example, the higher the contextual interpretation ability, the higher the accuracy of contextual interpretation may be. For example, an intermediate model may include a large language model having about 8B parameters.

[0090] A premium model may be a model that has a higher inference quality than an intermediate model. For example, the quality of the output data produced by a premium model may be higher than the quality of the output data produced by other models (e.g., a starter model, a cost-effective model, an intermediate model). For example, an intermediate model may include a large language model with about 70 billion parameters.

[0091] The models included in the exemplified multiple candidate models are for convenience of explanation only and should not be considered as limiting the embodiments thereto. As a non-limiting example, the multiple candidate models may include a model trained according to a specific type of emotional state and / or a model trained for a specific purpose.

[0092] According to one embodiment, the target model determination module (319) may include a determination model trained to determine a target model among a plurality of candidate models based on emotion information and / or model information (321). For example, the electronic device (101) may identify an input prompt that includes text corresponding to emotion information and text corresponding to model information (321). For example, the text corresponding to model information (321) may include list information for a plurality of candidate models. For example, the content of the input prompt may be expressed in natural language. For example, the electronic device (101) may obtain an output prompt representing the target model by providing the input prompt to the determination model. For example, the output prompt may include the value of a hyperparameter set in the target model. By using the determination model, the electronic device (101) can determine the target model without changing separate settings, even if the candidate models are expanded or changed.

[0093] According to one embodiment, candidate models may be referred to as candidate artificial intelligence models in that they include artificial intelligence models. For example, candidate models may include generative artificial intelligence models (e.g., language models, large language models, multimodal models). For example, a service based on a generative artificial intelligence model (e.g., chatbot, virtual assistant) may be provided by the electronic device (101). In the present disclosure, providing data (e.g., information about a voice signal, a prompt indicating a voice signal) to a generative artificial intelligence model may be understood to include providing data to a service based on a generative artificial intelligence model. That is, in the present disclosure, providing data (e.g., information about a voice signal, a prompt indicating a voice signal) to a model within the candidate models may include providing data to a generative artificial intelligence model and / or providing data to a service based on a generative artificial intelligence model.

[0094] For example, a generative AI model can be referred to as an artificial neural network-based model that has learned a large amount of data (e.g., text data, images, audio, video) through prior training. For example, a generative AI model may contain relatively more parameters (e.g., more than 10 billion) than existing general models. For example, a generative AI model may use a transformer artificial neural network structure based on an attention mechanism.

[0095] The attention mechanism is a technique that helps artificial intelligence models focus on important parts within input data. The attention mechanism can be utilized to predict output data by predicting the extent to which parts of time-series input data (e.g., input data such as voice or video, or input data for a specific layer of a neural network) contribute to the intermediate or final output of the neural network. While the recurrent neural network (RNN) structure, which processes each element of a sequence sequentially, suffers from degraded prediction performance when there is information dependency over long time-series distances, the attention mechanism can account for information dependency over long time-series distances by controlling the degree of weighted attention within the overall context (or part thereof) of the input data.

[0096] A transformer can be composed of an encoder-decoder structure. The encoder processes input data to output compressed information (e.g., contextual representation), and the decoder processes the compressed information to output data in token units. Each of the encoder and decoder may include an independent attention network and a cross-attention network connecting the encoder and decoder.

[0097] According to one embodiment, the training of a generative AI model may include pre-training and / or fine-tuning. Pre-training is a process of enabling the generative AI model to acquire general linguistic knowledge using a large amount of data (e.g., text data, images, audio, video), and may include, for example, self-supervised learning that predicts the next word using the previous sequence of words in a sequence of text. Fine-tuning is a process of training the generative AI model to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task, and the generative AI model may be further supervised (or adaptive) using a dataset suitable for the domain purpose based on the pre-trained model. The generative AI model may perform tasks using text input containing natural language called a prompt.

[0098] According to one embodiment, fine-tuning may be omitted during the training of a generative AI model. To enhance the performance of a task desired by the user, the prompts input to the generative AI model can be controlled. Examples of the task and / or guides for performing the task may be additionally provided in the prompts, such as in-context learning or zero-shot / few-shot learning. Examples of publicly available generative AI models include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer).

[0099] According to one embodiment, "inputting an input prompt into a generative artificial intelligence model" may mean "inputting an input prompt into an inference engine based on a generative artificial intelligence model." For example, "output of a generative artificial intelligence model for an input prompt" may mean output information of the last neural network layer of a generative artificial intelligence model (or output information modified through additional processing) obtained when an input prompt is input into an inference engine based on a generative artificial intelligence model.

[0100] The task execution module (327) may be used to perform one or more tasks according to a task plan output by the target model. A task may be referred to as a function and / or processor executed by at least one processor (e.g., at least one processor (200)) through a specified code. For example, a task may be referred to as a unit of operation of an electronic device (101) performed by at least one processor (200). For example, a task may be implemented as a function, an API (application programming interface), a deep link, a URL (Uniform Resource Locator), an intent, and / or a command, but the embodiments are not limited.

[0101] For example, the electronic device (101) may provide information about a voice signal to a target model. For example, the information about the voice signal may include a prompt containing text corresponding to the voice signal. For example, the electronic device (101) may obtain a task plan for providing a response to the voice signal by providing the prompt to the target model. For example, the task plan may include a function to be executed by the electronic device (101). For example, the electronic device (101) may perform one or more tasks according to the task plan by executing the function. As an example, but not limited to, the electronic device (101) may retrieve information necessary to perform each of the one or more tasks. The electronic device (101) may obtain, identify, or store the results of the execution of each of the one or more tasks. The electronic device (101) may obtain, identify, or store the results of the execution according to the function executed for each of the one or more tasks. The task execution module (327) may provide the response generation module (329) with the results of one or more tasks and / or the results of execution according to the above function. The task execution module (327) may be referred to as a task execution module.

[0102] A response generation module (329) may be used to obtain or generate response information. For example, the response generation module (329) may obtain response information based on emotional information, using the results of one or more tasks and / or the results of execution according to a function. For example, the response information may include a method of outputting a response to be provided by the electronic device (101). For example, the response information may include data representing content (e.g., image, text, audio) for providing (or expressing) a response. For example, the response information may include data representing characteristics (e.g., size, style, length, color) of the content for providing a response. For example, the response information may include data representing user interface (UI) elements related to the response and / or data representing animation elements related to the response. For example, the response generation module (329) may provide response information to the response providing module (331).

[0103] The response providing module (331) may be used to provide a response based on response information. For example, the response may be a response to a voice signal. For example, the electronic device (101) may output the response using the response providing module (331). For example, providing a response to a voice signal may include providing or executing a response function to a voice signal. For example, the electronic device (101) may output the response through a speaker (e.g., speaker (240)) by applying text-to-speech (TTS) to the response text represented by the response information. For example, the electronic device (101) may output the response through the speaker (240) by playing the response text as audio. For example, the response output through the speaker (240) may be output as a synthesized voice (e.g., TTS voice) according to the emotional information. For example, the synthesized voice based on emotional information may be a voice in which the tone of the voice is adaptively set (or adjusted) according to the emotional information. For example, the synthesized voice based on emotional information may be a voice in which the intonation of the voice is adaptively set (or adjusted) according to the emotional information. The electronic device (101) providing a response to a voice signal will be exemplified in FIG. 7, FIG. 8a, FIG. 8b, FIG. 9a, and / or FIG. 9b.

[0104] FIGS. 4a and 4b illustrate an example of an electronic device that determines a target model among a plurality of candidate models based on emotional information. The plurality of candidate models may be referred to as a plurality of artificial intelligence models. The target model may be referred to as a target artificial intelligence model. The target model may be referred to as the optimal model determined among the candidate models based on emotional information. For example, the target model may not be a separate model, but a model that is decided or selected to be used among the candidate models.

[0105] Referring to FIG. 4a, in example (401), the electronic device (101) can receive an audio input through a microphone (e.g., microphone (250)). The audio input may include a voice signal (411). For example, the voice signal (411) may include a voice such as "What is the weather like today?"

[0106] The electronic device (101) can obtain text information for a voice signal (411) using an ASR module (e.g., ASR module (303)). For example, the text information may include text such as 'What is the weather like today?'

[0107] The electronic device (101) can identify information about the tone of the voice signal (411) using a tone identification module (e.g., tone identification module (305)). For example, information about the tone of the voice signal (411) may indicate calmness.

[0108] The electronic device (101) can obtain information indicating an environment based on a background noise signal included in the audio input by using an environment identification module (e.g., environment identification module (307)). For example, an environment based on a background noise signal may indicate an indoor environment (e.g., inside a house).

[0109] The electronic device (101) can obtain emotion estimation information based on user information by using a user information collection module (e.g., user information collection module (309)) and / or an emotion estimation information identification module (e.g., emotion estimation information identification module (311)). For example, the user information may be user information collected within a time interval defined by the reception of audio input. For example, the user information may indicate that the user has no direct interaction with the electronic device (101).

[0110] The electronic device (101) can acquire sensing information by using a sensing information management module (e.g., a sensing information management module (315)). For example, the sensing information can be acquired using at least one sensor within the wearable device (313). For example, a communication link can be established between the wearable device (313) and the electronic device (101). For example, the wearable device (313) can transmit the sensing information to the electronic device (101) via the communication link using the communication circuit of the wearable device (313). For example, the electronic device (101) can acquire the sensing information. For example, the sensing information may indicate that the heart rate is within a reference range.

[0111] The electronic device (101) can identify emotional information based on text information regarding the voice signal (411), information regarding the tone of the voice signal (411), information indicating the environment based on the background noise signal, emotional estimation information based on user information, and / or sensing information. For example, the emotional information of the electronic device (101) may correspond to a first emotional type (e.g., calmness). For example, the electronic device (101) can determine a first target model (421) among candidate models (420) using a target model determination module (e.g., target model determination module (319)) based on the emotional information. For example, the first target model (421) may be a cost-effective model. For example, the first target model (421) may perform routine tasks faster than other models. The electronic device (101) may provide information regarding the voice signal (411) to the first target model (421). For example, information regarding the voice signal (411) may include a prompt containing text representing the voice signal (411). The electronic device (101) may provide a response to the voice signal (411) using the first target model (421). For example, the electronic device (101) may provide a response indicating today's weather based on location and / or time. For example, the electronic device (101) may provide a response indicating the maximum temperature, minimum temperature, fine dust concentration, and / or ultrafine dust concentration. As an example, but not limited to, the electronic device (101) may provide weather for said area based on the user identifying the area where they have a schedule. For example, the response to the voice signal (411) may be displayed in full screen via a display (e.g., display (220)). For example, the full screen above states, "Today's weather is clear, and the temperature in Seoul is a high of 18℃ and a low of 10℃. Fine dust is 'moderate', and ultrafine dust is 'good'.It can include text such as, "Suwon, where you have an event at 3 PM, is expected to have a high of 15℃ and may rain, so please bring an umbrella."

[0112] In example (402), the electronic device (101) can receive an audio input through a microphone (250). The audio input may include a voice signal (412). For example, the voice signal (412) may include a voice such as "What is the weather like today?!!"

[0113] The electronic device (101) can obtain text information for a voice signal (412) using an ASR module (303). For example, the text information may include text such as 'How is the weather today?!!'

[0114] The electronic device (101) can identify information about the tone of the voice signal (412) using a tone identification module (e.g., tone identification module (305)). For example, information about the tone of the voice signal (412) may indicate anger and / or discomfort.

[0115] The electronic device (101) can obtain information indicating an environment based on a background noise signal included in the audio input by using an environment identification module (e.g., environment identification module (307)). For example, an environment based on a background noise signal may indicate a relatively noisy environment (e.g., a busy street, a performance venue).

[0116] The electronic device (101) can obtain emotion estimation information based on user information by using a user information collection module (e.g., user information collection module (309)) and / or an emotion estimation information identification module (e.g., emotion estimation information identification module (311)). For example, the user information may be user information collected within a time interval defined by the reception of audio input. For example, the user information may represent a user who is nervously searching the internet.

[0117] The electronic device (101) can acquire sensing information by using a sensing information management module (e.g., a sensing information management module (315)). For example, the sensing information may indicate that the heart rate is higher than a reference range and / or that the user is walking.

[0118] The electronic device (101) can identify emotional information based on text information regarding the voice signal (412), information regarding the tone of the voice signal (412), information indicating the environment based on the background noise signal, emotional estimation information based on user information, and / or sensing information. For example, the emotional information of the electronic device (101) may correspond to a second emotional type (e.g., high stress). For example, the electronic device (101) can determine a second target model (422) among candidate models (420) using a target model determination module (e.g., target model determination module (319)) based on the emotional information. For example, the second target model (422) may be a starter model. For example, the second target model (422) may be a model trained for the performance of a relatively fast task. The electronic device (101) can provide a response to the voice signal (412) using the second target model (422). For example, the electronic device (101) may provide a response containing relatively important information regarding today's weather. For example, the response to the voice signal (412) may be displayed using a UI object via a display (e.g., display (220)). For example, the UI object may be displayed along the top of the display (220). For example, the UI object may include text such as "Sunny, high 18°C, low 10°C expected."

[0119] Referring to FIG. 4b, in example (403), the electronic device (101) can receive an audio input through a microphone (250). The audio input may include a voice signal (413). For example, the voice signal (413) may include a voice such as "What is the weather like today?..."

[0120] The electronic device (101) can obtain text information for a voice signal (413) using an ASR module (303). For example, the text information may include text such as 'What is the weather like today?...'.

[0121] The electronic device (101) can identify information about the tone of the voice signal (413) using a tone identification module (e.g., tone identification module (305)). For example, information about the tone of the voice signal (413) may indicate sadness.

[0122] The electronic device (101) can obtain information indicating an environment based on a background noise signal included in the audio input by using an environment identification module (e.g., environment identification module (307)). For example, an environment based on a background noise signal may indicate a relatively quiet environment (e.g., a library, a classroom).

[0123] The electronic device (101) can obtain emotion estimation information based on user information by using a user information collection module (e.g., user information collection module (309)) and / or an emotion estimation information identification module (e.g., emotion estimation information identification module (311)). For example, the user information may be user information collected within a time interval defined by the reception of audio input. For example, the user information may indicate scrolling through a social network service (SNS) application.

[0124] The electronic device (101) can acquire sensing information by using a sensing information management module (e.g., a sensing information management module (315)). For example, the sensing information may indicate that the heart rate is lower than a reference range.

[0125] The electronic device (101) can identify emotional information based on text information regarding a voice signal (413), information regarding the tone of the voice signal (413), information indicating the environment based on a background noise signal, emotional estimation information based on user information, and / or sensing information. For example, the emotional information of the electronic device (101) may correspond to a third emotional type (e.g., sadness). For example, the electronic device (101) can determine a third target model (423) among candidate models (420) using a target model determination module (e.g., target model determination module (319)) based on the emotional information. For example, the third target model (423) may be a premium model. For example, the third target model (423) may be a model with a relatively slow inference speed but relatively high inference quality. For example, the electronic device (101) can perform a task to provide a response to a voice signal (413) using a third target model (423). The electronic device (101) can use a task execution module (e.g., task execution module (327)) to retrieve information necessary for performing the task. For example, the necessary information may identify a message indicating a rejection of employment, the user's schedule, the location of the electronic device (101), and / or a point of interest (POI). The electronic device (101) can provide a response to a voice signal (413) using the third target model (423). For example, the electronic device (101) can provide a response by outputting a relatively calm voice through a speaker (e.g., speaker (240)). For example, the response may include relatively calm music. For example, the above response is, "It is sunny today, so it is a good day for a light walk. Seoul is expected to reach 18℃ at its warmest and 10℃ at its coldest, so I think you will be able to spend the day comfortably outdoors with just a long-sleeved shirt, a cardigan, or a hat."It can include a voice like, "If you don't have any particular plans, why not take a leisurely walk in a nearby park? If you have a cup of your favorite popular Flatccino at the newly opened cafe next door, I think it could make for a day filled with pleasant memories. Please feel free to let me know whenever you are free, and I will be listening attentively."

[0126] In example (404), the electronic device (101) can receive an audio input through a microphone (250). The audio input may include a voice signal (414). For example, the voice signal (414) may include a voice such as "What is the weather like today?"

[0127] The electronic device (101) can obtain text information for a voice signal (414) using an ASR module (303). For example, the text information may include text such as 'What is the weather like today?'

[0128] The electronic device (101) can identify information about the tone of the voice signal (414) using a tone identification module (e.g., tone identification module (305)). For example, information about the tone of the voice signal (414) can indicate pleasure.

[0129] The electronic device (101) can obtain information indicating an environment based on a background noise signal included in the audio input by using an environment identification module (e.g., environment identification module (307)). For example, an environment based on a background noise signal may indicate an environment containing some noise.

[0130] The electronic device (101) can obtain emotion estimation information based on user information by using a user information collection module (e.g., user information collection module (309)) and / or an emotion estimation information identification module (e.g., emotion estimation information identification module (311)). For example, the user information may be user information collected within a time interval defined by the reception of audio input. For example, the user information may indicate that a short-form video is being played and / or that the degree of movement of the electronic device (101) is relatively low.

[0131] The electronic device (101) can acquire sensing information by using a sensing information management module (e.g., a sensing information management module (315)). For example, the sensing information may indicate that the heart rate is within a reference range.

[0132] The electronic device (101) can identify emotional information based on text information regarding a voice signal (414), information regarding the tone of the voice signal (414), information indicating the environment based on a background noise signal, emotional estimation information based on user information, and / or sensing information. For example, the emotional information of the electronic device (101) may correspond to a fourth emotional type (e.g., joy). For example, the electronic device (101) can determine a fourth target model (424) among candidate models (420) using a target model determination module (e.g., target model determination module (319)) based on the emotional information. For example, the fourth target model (424) may be an intermediate model learned based on the fourth emotional type. For example, the fourth target model (424) may be a model in which the value of a hyperparameter indicating the creativity of the artificial intelligence model is set relatively high. For example, the electronic device (101) may search for or obtain internet news information to perform inference using the fourth target model (424). For example, the electronic device (101) may provide a response including at least one method for a voice signal using the fourth target model (424). For example, the electronic device (101) may provide a response by outputting the voice through a speaker (e.g., speaker (240)). For example, the response may include a voice such as, "Today is sunny with a high of 18°C ​​and a low of about 10°C. I think a light cardigan will be enough for the day! If you don't have any particular plans, it would be good to contact friends A and B for the first time in a while, try playing a new game, or watch a movie!"

[0133] FIG. 5 illustrates an example of an electronic device that changes a target model based on a change in emotional information. The target model may be referred to as a target artificial intelligence model.

[0134] Referring to FIG. 5, in example (501), the electronic device (101) may receive an audio input through a microphone (e.g., microphone (250)). For example, the audio input may include a first voice signal (e.g., "How is the weather?"). Since example (501) may be substantially the same as example (401) of FIG. 4a, redundant content is omitted. For example, text information may include text such as 'How is the weather?'. For example, information regarding the tone of the first voice signal may indicate calmness. For example, the environment according to the background noise signal may indicate an indoor environment (e.g., inside a house). For example, the user information may indicate that the user is not interacting directly with the electronic device (101). For example, sensing information may indicate that the heart rate is within a reference range.

[0135] The electronic device (101) can identify emotional information. For example, the emotional information of the electronic device (101) may correspond to a first emotional type (e.g., calmness). For example, the electronic device (101) can determine a first target model (521) (e.g., first target model (421)) from among candidate models (520) (e.g., candidate models (420)) by using a target model determination module (e.g., target model determination module (319)) based on the emotional information. For example, the candidate models (520) may be referred to as candidate artificial intelligence models. For example, the first target model (521) may be a cost-effective model. The electronic device (101) can provide a response to a first voice signal using the first target model (521). For example, the electronic device (101) may provide a response to the first voice signal by outputting voice through a speaker (e.g., speaker (240)). For example, the response to the first voice signal may include a voice such as, "The weather today is clear, and the temperature in Seoul is 18°C ​​at its highest and 10°C at its lowest. Fine dust levels are moderate, and ultrafine dust levels are 'good'. Suwon, where you have an event at 3 PM, is expected to have a high of 15°C, and it may rain, so please bring an umbrella."

[0136] In example (502), the electronic device (101) may receive an audio input containing a second voice signal (e.g., "How's the weather!!") through the microphone (250), even though the electronic device (101) has provided a response to the first voice signal. Since example (502) may be substantially identical to example (402) of FIG. 4a, redundant content is omitted. For example, text information may include text such as "How's the weather!!". For example, information regarding the tone of the second voice signal may indicate anger and / or discomfort. For example, the environment based on the background noise signal may indicate an indoor environment (e.g., inside a house). For example, the user information may include data indicating a relatively strong shaking of the electronic device (101). For example, the user information may include previous emotion information. For example, the previous emotion information may include emotion information identified in example (501).

[0137] The electronic device (101) can identify emotional information. For example, the emotional information of the electronic device (101) can correspond to a second emotional type (e.g., high stress, anger). For example, based on the emotional information, the electronic device (101) can determine a second target model (522) (e.g., second target model (422)) from among candidate models (520) (e.g., candidate models (420)) by using a target model determination module (e.g., target model determination module (319)). For example, the second target model (522) may be an intermediate model or a premium model. The electronic device (101) can provide a response to a second voice signal using the second target model (522). For example, the electronic device (101) can provide a response to a second voice signal by outputting the voice through a speaker (e.g., speaker (240)). For example, a response to the second voice signal could be a voice like, "I apologize. I just told you that today's weather is clear, with temperatures in Seoul ranging from a high of 18°C ​​to a low of 10°C. Additional information I can provide includes the news forecast predicting increased traffic, fine dust levels, ultrafine dust levels, this week's and next week's weather, or the weather in your specific area. I would appreciate it if you could provide a little more detail so I can tell you the specific items you are looking for," and / or, "I apologize. To verify the details further, I have compiled the results from local social media and various weather sites. It is reported that today is clear everywhere, with the highest temperatures not exceeding 19°C and the coldest temperatures not dropping below 9°C. Although there is a less than 10% chance of rain, it is expected to be around 1mm, so you likely won't need to bring an umbrella. It will get colder after 4 PM, when you are meeting your friend, so if you plan to stay out late, I recommend dressing warmly. Please let me know if you have any further questions, and I will look into it and get back to you." It can include voice.

[0138] The electronic device (101) can determine a second target model (522) among candidate models (520) based on the difference between the first voice signal of example (501) and the second voice signal of example (502). For example, the electronic device (101) can identify that the user is not satisfied with the response provided in example (501). For example, the electronic device (101) can determine a second target model (522) among candidate models (520) based on the emotional information identified in example (501) and the emotional information identified in example (502). For example, the electronic device (101) may include emotional information including changes in the user's emotions. The electronic device (101) can adaptively determine a target model for inference regarding the voice signal according to the emotional information. For example, the electronic device (101) may determine the second target model (522) among the first target model (421) and the second target model (522) for performing inference according to the user's emotional change. The size of the second target model (522) may be larger than the size of the first target model (521). For example, a larger model size may include a larger number of model parameters. For example, the number of parameters of the second target model (522) may be greater than the number of parameters of the first target model (521).

[0139] For example, the inference quality of the second target model (522) may be higher than the inference quality of the first target model (521). For example, the length of the response to the second voice signal may be longer than the length of the response to the first voice signal.

[0140] FIGS. 6a and 6b illustrate examples of operations of an electronic device (e.g., electronic device (101)) that provides a response to a voice signal. In FIGS. 6a and 6b, a model (e.g., candidate models, at least one target model) may be understood to include an artificial intelligence model. For example, the artificial intelligence model may include a generative artificial intelligence model.

[0141] Referring to FIG. 6a, in operation 601, an electronic device (101) (e.g., at least one processor (200)) may receive an audio input through a microphone (e.g., a microphone (250)). For example, the audio input may include a voice signal of a user of the electronic device (101) (e.g., voice signal (411), voice signal (412), voice signal (413), voice signal (414)) and / or a background noise signal. For example, the voice signal may correspond to a user's speech. For example, the voice signal may represent a user's voice command. For example, the background noise signal may correspond to a background sound. For example, the background noise signal may represent information about the environment obtained through the microphone (250) via the audio input.

[0142] In operation 603, an electronic device (101) (e.g., at least one processor (200)) may obtain information regarding the tone of a voice signal included in an audio input and user information regarding the reception of the audio input. For example, information regarding the tone of the voice signal may indicate acoustic characteristics of the voice signal. For example, information regarding the tone of the voice signal may include the pitch of the voice signal, the speed of the voice signal, and / or the intensity of the voice signal. For information regarding the tone of the voice signal, the descriptions of the tone identification module (305) of FIG. 3 may be referenced.

[0143] User information resulting from the reception of audio input may include user information collected within a time interval defined by the reception of the audio input. For example, for user information resulting from the reception of audio input, descriptions of the user information collection module (309) and / or emotion estimation information identification module (311) of FIG. 3 may be referenced.

[0144] In operation 605, an electronic device (101) (e.g., at least one processor (200)) can identify emotional information based on information about the tone of a voice signal and user information. For emotional information, the descriptions of the emotional information identification module (317) of FIG. 3 may be referenced.

[0145] According to one embodiment, information regarding the tone of a voice signal may be information for which correction processing has been performed according to information regarding the environment in which the audio input was obtained. For example, an electronic device (101) may obtain first information regarding the tone of a voice signal from an audio input. For example, the electronic device (101) may perform correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal. For example, the electronic device (101) may obtain second information regarding the tone of the voice signal by performing correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal. For example, the electronic device (101) may identify emotional information based on the second information regarding the tone of the voice signal and user information. For example, by performing correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal, the quality of the emotional information identified in operation 605 may be relatively high. For example, the emotional information identified in operation 605 may be substantially the same as the actual emotional information of the user of the electronic device (101).

[0146] In operation 607, the electronic device (101) (e.g., at least one processor (200)) can determine at least one target model among candidate models based on emotional information. For example, for operation 607, the descriptions of the target model determination module (319) of FIG. 3 may be referenced.

[0147] In operation 609, an electronic device (101) (e.g., at least one processor (200)) may provide a response to a voice signal based on providing information about the voice signal to at least one target model. For operation 609, descriptions of the response generation module (329) and / or response providing module (331) of FIG. 3 may be referenced.

[0148] According to one embodiment, as an example but not limited to, the electronic device (101) may provide emotional information to at least one target model. For example, the electronic device (101) may provide information regarding emotional information and / or voice signals to at least one target model. For example, the electronic device (101) may generate or obtain a prompt including text representing a voice signal and / or text representing emotional information. For example, the electronic device (101) may generate or obtain a response to the voice signal by providing the prompt to at least one target model. For example, the electronic device (101) may provide a response in a different way depending on the emotional information by providing the prompt to at least one target model.

[0149] According to one embodiment, information regarding a voice signal may include or include a prompt containing text representing the voice signal. For example, the text may be generated by applying speech-to-text (STT) to the voice signal. For example, the prompt may be generated in a manner according to the at least one target model. For example, the format of the prompt may be a format corresponding to the at least one target model. For example, the format of the prompt may be a format optimized for the at least one target model. For example, the format optimized for the at least one target model may be determined according to the manner in which the at least one target model was learned. For example, the format optimized for the at least one target model may be determined according to the content learned by the at least one target model.

[0150] According to one embodiment, the electronic device (101) may provide a response to a voice signal in a manner according to emotional information. For example, the electronic device (101) providing a response in a different manner according to emotional information will be explained with reference to FIGS. 7 through 9b.

[0151] As an example not limited to, the electronic device (101) may further provide information about the environment in which the audio input was acquired to at least one target model. For example, the electronic device (101) may provide information about the environment and information about the voice signal to at least one target model. For example, the electronic device (101) may provide a response to the voice signal by providing information about the environment and information about the voice signal to at least one target model. For example, by providing information about the environment to at least one target model, the inference ability of at least one target model may be enhanced. For example, the quality of the first response to the voice signal acquired according to the information about the environment and information about the voice signal may be higher than the quality of the second response to the voice signal acquired according to the information about the voice signal.

[0152] Referring to FIG. 6b, operations 611, 613, and / or 615 may represent specific embodiments of operation 609 of FIG. 6a.

[0153] In operation 611, the electronic device (101) (e.g., at least one processor (200)) can obtain information indicating at least one function by providing information about a voice signal to at least one target model. For example, the information indicating at least one function may further indicate a task corresponding to at least one function. For example, the task may be performed to provide a response to the voice signal. For example, at least one function may be referred to as a function for performing the task. For example, the electronic device (101) may perform the task by executing at least one function.

[0154] For example, at least one target model may include at least one generative artificial intelligence model. For example, information representing at least one function may include a prompt representing at least one function.

[0155] As an example not limited to, the prompt may include an output method based on emotion information. For example, the output method may relate to the output format of the response and / or the length of the response and the data size of the response. For example, the output method may include a method for specifying the length of the response displayed as text (e.g., 50 words or less). For example, the data size of the response indicated by the output method when the emotion information corresponds to anger may be smaller than the data size of the response indicated by the output method when the emotion information corresponds to sadness. However, the embodiments are not limited.

[0156] In operation 613, the electronic device (101) (e.g., at least one processor (200)) can obtain response information by using the result of execution of at least one function. For operation 613, the descriptions of the task execution module (327) and / or response generation module (329) of FIG. 3 may be referenced.

[0157] In operation 615, an electronic device (101) (e.g., at least one processor (200)) may provide a response based on at least one of response information or emotion information. For operation 615, the descriptions of the response providing module (331) of FIG. 3 may be referenced.

[0158] FIG. 7 illustrates an example of an electronic device that adaptively outputs a response to a voice signal according to emotional information. For example, the electronic device (101) may provide or output a response to a voice signal in a different way according to emotional information in operation 609 of FIG. 6a. For example, the electronic device (101) may adaptively output a response to a voice signal according to emotional information based on providing emotional information to at least one target model. For example, the target model may be referred to as a target artificial intelligence model.

[0159] Referring to FIG. 7, in example (701), the electronic device (101) can identify emotional information corresponding to a first emotional type (e.g., calmness). Since example (701) can substantially correspond to example (401) of FIG. 4a, redundant content is omitted. The electronic device (101) can provide a response to a voice signal based on the emotional information. For example, the electronic device (101) can display the response to the voice signal on an execution screen (710) via a display (220). The execution screen (710) may include a visual object (711), a visual object (712), and / or a visual object (713). For example, the visual object (711) may correspond to the voice signal. For example, the visual object (711) may be obtained by applying STT to the voice signal. For example, text included in a visual object (711) (e.g., 'Today's weather') can correspond to a voice signal.

[0160] According to one embodiment, a visual object (712) may be displayed through a display (220) below a visual object (711). The visual object (712) may include a response to a voice signal. For example, text included in the visual object (712) (e.g., 'The weather in B-dong, A-gu today is clear.') may be a response to a voice signal.

[0161] According to one embodiment, a visual object (713) may be displayed through a display (220) below a visual object (712). For example, the visual object (713) may include content (e.g., a weather image) that represents a response to a voice signal. For example, the visual object (713) may indicate the temperature as well as the concentration of fine dust and / or ultrafine dust.

[0162] In example (702), the electronic device (101) can identify emotional information corresponding to a second type of emotion (e.g., high stress, impatience). Since example (702) can substantially correspond to example (402) of FIG. 4a, redundant content is omitted. For example, in example (702), it can be assumed that the user is late for an appointment. For example, while running a navigation application, the electronic device (101) can acquire a voice signal (e.g., "Today's weather") through a microphone (e.g., microphone (250)). The electronic device (101) can provide a response to the voice signal based on the emotional information. For example, the electronic device (101) can display the response to the voice signal as a UI object (721) through the display (220).

[0163] According to one embodiment, the electronic device (101) may display a UI object (721) superimposed on an execution screen (720). For example, the execution screen (720) may be referred to as a screen corresponding to the execution of a navigation application. For example, the electronic device (101) may display the UI object (721) along the top of a display (220). For example, the electronic device (101) may display the UI object (721) along the top of an execution screen (720). The UI object (721) may include a response to a voice signal. For example, text included in the UI object (721) (e.g., 'Secretary: Clear, high 15 degrees low 3 degrees') may be a response to a voice signal.

[0164] In example (701), the electronic device (101) can provide a response to a voice signal in relatively detailed terms based on a first emotion type. In example (702), the electronic device (101) can provide a response to a voice signal in relatively compressed terms based on a second emotion type. For example, the size of the response data provided in example (701) may be larger than the size of the response data provided in example (702). For example, the amount of information in the response provided in example (701) may be greater than the amount of information in the response provided in example (702). The electronic device (101) can adaptively adjust the method of the response to the voice signal according to the emotion information.

[0165] FIGS. 8a and 8b illustrate examples of electronic devices that adaptively control an external electronic device according to emotional information. For example, in operation 609 of FIG. 6a, the electronic device (101) may provide or output a response to a voice signal in a different way according to the emotional information. For example, the electronic device (101) may adaptively output a response to a voice signal according to the emotional information based on providing emotional information to at least one target model. For example, the target model may be referred to as a target artificial intelligence model.

[0166] Referring to FIG. 8a, the electronic device (101) can receive an audio input containing a voice signal (811) (e.g., "It's hot") through a microphone (e.g., microphone (250)). FIG. 8a may correspond to the example (401) of FIG. 4a. Redundant content is omitted. The electronic device (101) can identify emotional information. For example, the emotional information may correspond to a first emotional type (e.g., calmness). For example, based on the emotional information, the electronic device (101) can determine a first target model (e.g., first target model (421)) among candidate models (e.g., candidate models (420)) using a target model determination module (e.g., target model determination module (319)). For example, the candidate models (420) may be referred to as candidate artificial intelligence models. The electronic device (101) may provide information (e.g., a prompt) regarding a voice signal (811) to a first target model. The electronic device (101) may identify an external electronic device (801) (e.g., an air conditioner) for executing a function corresponding to the voice signal (411) among at least one external electronic device registered to the electronic device (101). For example, the function corresponding to the voice signal (411) may be related to a cooling function. For example, the electronic device (101) may adaptively determine a setting value for the cooling function according to emotional information. For example, the electronic device (101) may transmit a control signal for executing the function to the external electronic device (801) via a communication circuit (e.g., a communication circuit (230)). For example, the control signal may include a first set temperature (813) and / or a first set wind speed (815). The above control signal can cause the external electronic device (801) to activate the external electronic device (801) by setting the set temperature of the external electronic device (801) to a default set temperature.

[0167] Referring to FIG. 8b, the electronic device (101) can receive an audio input containing a voice signal (821) (e.g., "It's hot!!!") through a microphone (e.g., microphone (250)). FIG. 8b may correspond to the example (402) of FIG. 4a. Redundant content is omitted. The electronic device (101) can identify emotional information. For example, the emotional information may correspond to a second emotional type (e.g., high stress, anger). For example, based on the emotional information, the electronic device (101) can determine a second target model (e.g., second target model (422)) from among candidate models (e.g., candidate models (420)) using a target model determination module (e.g., target model determination module (319)). The electronic device (101) can provide information (e.g., a prompt) regarding the voice signal (821) to the second target model. The electronic device (101) can identify an external electronic device (801) (e.g., an air conditioner) for executing a function corresponding to a voice signal (821) among at least one external electronic device registered to the electronic device (101). For example, the function corresponding to the voice signal (821) may be related to a cooling function. For example, the electronic device (101) may adaptively determine a setting value for the cooling function according to emotional information. For example, the electronic device (101) may transmit a control signal for executing the function to the external electronic device (801) through a communication circuit (e.g., a communication circuit (230)). For example, the control signal may include a second set temperature (823) and / or a second set wind speed (825). For example, the temperature indicated by the second set temperature (823) may be lower than the temperature indicated by the first set temperature (813) of FIG. 8A. For example, the second set wind speed (825) may be higher than the first set wind speed (815) of FIG. 8a.As a non-limiting example, the control signal may cause the external electronic device (801) to activate the external electronic device (801) by maintaining the set temperature of the external electronic device (801) at the lowest set temperature for a certain period of time (e.g., 5 minutes).

[0168] FIGS. 9a and 9b illustrate examples of electronic devices that adaptively determine content based on emotional information. For example, the electronic device (101) may provide or output a response to a voice signal in a different way based on emotional information in operation 609 of FIG. 6a. For example, the electronic device (101) may adaptively output a response to a voice signal based on emotional information, based on providing emotional information to at least one target model. For example, the at least one target model may be referred to as at least one target artificial intelligence model.

[0169] Referring to FIG. 9a, the electronic device (101) can receive an audio input including a voice signal (910) (e.g., "Play music") through a microphone (e.g., microphone (250)). FIG. 9a may correspond to the example (401) of FIG. 4a. Redundant content is omitted. The electronic device (101) can identify emotional information. For example, the emotional information may correspond to a first emotional type (e.g., calmness). For example, based on the emotional information, the electronic device (101) can determine a first target model (e.g., first target model (421)) from among candidate models (e.g., candidate models (420)) using a target model determination module (e.g., target model determination module (319)). For example, the candidate models (420) may be referred to as candidate artificial intelligence models. The electronic device (101) can provide information (e.g., a prompt) about the voice signal (811) to the first target model.

[0170] The electronic device (101) can output a first audio through a speaker (240) using a first target model. By outputting the first audio through the speaker (240), the electronic device (101) can provide a response to a voice signal (910). For example, the electronic device (101) can adaptively determine the first audio based on emotional information. For example, the first audio may be referred to as relatively exciting music. For example, the first audio may include the latest songs. For example, the electronic device (101) can run a music application. For example, the electronic device (101) can display an execution screen (915) for the music application through a display (220). For example, the execution screen (915) may include an image (e.g., an album cover) corresponding to the first audio.

[0171] Referring to FIG. 9b, the electronic device (101) can receive an audio input containing a voice signal (920) (e.g., "Play music...") through a microphone (e.g., microphone (250)). FIG. 9b may correspond to the example (403) of FIG. 4b. Redundant content is omitted. The electronic device (101) can identify emotional information. For example, the emotional information may correspond to a second emotion type (e.g., sadness). For example, based on the emotional information, the electronic device (101) can determine a second target model (e.g., a third target model (423)) from among candidate models (e.g., candidate models (420)) using a target model determination module (e.g., target model determination module (319)). The electronic device (101) can provide information (e.g., a prompt) regarding the voice signal (821) to the second target model.

[0172] The electronic device (101) can output a second audio through a speaker (240) using a second target model. By outputting the second audio through the speaker (240), the electronic device (101) can provide a response to a voice signal (920). For example, the electronic device (101) can adaptively determine the second audio based on emotional information. For example, the second audio may be referred to as relatively calm music. For example, the electronic device (101) can identify the second audio by searching for audio that is pleasant to listen to in the case of the second emotional type. For example, the second audio may be audio that was previously identified with emotional information corresponding to the second emotional type and played by a user of the electronic device (101). For example, the second audio may include classical music. As an example not limited to, the pitch of the second audio in FIG. 9b may be lower than the pitch of the first audio in FIG. 9a. The tempo of the second audio in FIG. 9b may be slower than the tempo of the first audio in FIG. 9a. Additionally, the electronic device (101) may provide a response to the voice signal (920) by displaying a UI object (927) through the display (220). For example, the UI object (927) may include text to comfort the user. For example, the UI object (927) may include text such as "Cheer up. I'll play some music to cheer you up." For example, the electronic device (101) may display an execution screen (925) through the display (220) that includes an image (e.g., an album cover) corresponding to the UI object (927) and / or the second audio.

[0173] FIG. 10 is a schematic diagram of an exemplary artificial intelligence (AI) system. Each of the artificial intelligence models (e.g., candidate models) exemplified in this document may include at least a part of the AI ​​system (1000) of FIG. 10.

[0174] Referring to FIG. 10, the AI ​​system (1000) may include an input / output interface (1010), an AI framework (1020), a generative AI model (1030), and / or a knowledge repository (1090).

[0175] The input / output interface (1010) can receive input. The input may include user input and / or data obtained or generated by an electronic device (e.g., electronic device (101), electronic device (104), server (108)). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (200) or processor (120)), such as: illuminance data around the electronic device obtained from a sensor or sensor hub (e.g., auxiliary processor (123); attitude data (or orientation data) of the electronic device; temperature inside the electronic device (e.g., display (220)); or temperature of at least one processor (200); size information of the display area of ​​the display (220); and / or images obtained through an image sensor of the electronic device (e.g., included in a camera module (180)). The user input may include natural language, touch data obtained through a touch circuit included within the display panel (e.g., used to identify input from a finger and / or stylus), an image displayed (and / or to be displayed) on the display panel, and / or video. By example, without limitation, the user input may be received by an input / output interface (1010) along with context information. The context information may be described as additional information obtained in relation to the user input. The context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the context information may include information about one or more software applications executed within the electronic device at the time the user input is received.For example, the above situation information may include information regarding the location of the electronic device (or the location of the user of the electronic device) when the user input is received. For example, the user input may be integrated with the situation information. For example, the user input with the situation information integrated as the input may be received by the input / output interface (1010).

[0176] The input / output interface (1010) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI ​​system (1000) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.

[0177] The input / output interface (1010) can be described as a user question / response interface (1010).

[0178] The AI ​​framework (1020) can be used to obtain information (or data) about the input from the input / output interface (1010) and to control one or more components related to the AI ​​system (1000) using the obtained information.

[0179] For example, a prompt design component (1021) within an AI framework (1020) can generate or obtain prompts for a generative AI model (1030) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (1021) may be described as an AI component that uses a learning algorithm and / or a neural network to provide prompts that are enhanced over time. For example, the prompt design component (1021) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (1090)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (1030) (e.g., including an LLM or LMM).

[0180] For example, an API / plugin management component (1022) within the AI ​​framework (1020) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (1030). For example, the API / plugin management component (1022) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (1090)). For example, the API / plugin management component (1022) may support access to at least some of the data sources. For example, the API / plugin management component (1022) may be used to request another component (e.g., application / service component (1080)) that performs feedback (or response) according to the prompt. As an example without limitation, information obtained (or generated) through the API / plugin management component (1022) may be provided to the prompt design component (1021) for generating a prompt. As an example without limitation, information obtained (or generated) through the API / plugin management component (1022) may be provided to the generative AI model (1030).

[0181] For example, an improvement component (1023) within the AI ​​framework (1020) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1030). For example, the improvement component (1023) can determine or verify whether the content obtained from the generative AI model (1030) is related to the input. For example, the improvement component (1023) can determine or verify whether the content obtained from the generative AI model (1030) contains biased content. For example, the improvement component (1023) can determine or verify whether the content obtained from the generative AI model (1030) contains harmful content. For example, the improvement component (1023) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1030). For example, the improvement component (1023) may support providing a hint to the user to improve the content.

[0182] A generative AI model (1030) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information relative to the prompt, but relative to the prompt. For example, the feedback may include new content relative to the prompt. For example, the generative AI model (1030) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative AI model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, the generative AI model (1030) may include an LMM that generates the feedback by recognizing text, images, and / or speech.

[0183] As an example without limitation, the AI ​​framework (1020) and / or generative AI model (1030) may be included within an AI module (e.g., including a processing circuit) within the electronic device (101). For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (101) (e.g., at least one processor (200) or processor (120)). For example, the AI ​​module may be operatively coupled with a display driving circuit of the electronic device. For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0184] In an embodiment according to the present disclosure, an electronic device (e.g., electronic device (101)) may provide a response to a voice signal. For example, the electronic device (101) may adaptively determine a target model based on emotional information. For example, the electronic device (101) may adaptively determine the output method of the response based on emotional information. For example, because the electronic device (101) provides a response based on emotional information, it may provide a different type of response even if the text corresponding to the voice signal is substantially the same. Because the electronic device (101) adaptively provides a response based on emotional information, the electronic device (101) may evoke a friendly and human feeling in the user. The user experience of the electronic device (101) may be enhanced.

[0185] The electronic device (101) can enhance cost efficiency related to the model by determining a target model from among multiple candidate models based on emotional information, rather than consistently using a specific model. Since the electronic device (101) determines the target model based on emotional information, it can enhance cost efficiency related to the model while satisfying the user's satisfaction of the electronic device (101).

[0186] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0187] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0188] As described above, the electronic device may include a microphone. The electronic device may include a memory comprising one or more storage media for storing instructions. The electronic device may include at least one processor comprising processing circuitry. The instructions may cause the electronic device to receive an audio input through the microphone when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain information regarding the tone of a voice signal included in the audio input and user information collected within a time interval defined by the reception of the audio input when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to identify emotion information based on the information regarding the tone of the voice signal and the user information when executed individually or collectively by the at least one processor. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to determine at least one target artificial intelligence model among a plurality of candidate artificial intelligence models based on the emotion information. The above instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to provide a response to the voice signal based on providing information about the voice signal to the at least one target artificial intelligence model.

[0189] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. The instructions may cause the electronic device to obtain a prompt indicating at least one function by providing the information regarding the voice signal to the at least one generative artificial intelligence model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain response information by using the execution result of the execution of the at least one function when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to provide the response by processing the response information in a manner according to the emotion information when executed individually or collectively by the at least one processor.

[0190] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. When the instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to obtain a prompt including an output method according to the emotion information by providing the information regarding the voice signal to the at least one generative artificial intelligence model. When the instructions are executed individually or collectively by the at least one processor, they may cause the electronic device to provide the response having a data size according to the output method by using the prompt.

[0191] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. The information regarding the voice signal may include text representing the voice signal and may include a prompt generated in a manner according to the at least one generative artificial intelligence model.

[0192] According to one embodiment, the electronic device may further include a speaker. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to output the response set to a first tone through the speaker based on identifying the emotion information corresponding to a first emotion type. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to output the response set to a second tone different from the first tone through the speaker based on identifying the emotion information corresponding to a second emotion type different from the first emotion type.

[0193] According to one embodiment, the electronic device may further include a display. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display the response as an execution screen through the display based on identifying the emotion information corresponding to a first emotion type. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display the response as a user interface (UI) object through the display based on identifying the emotion information corresponding to a second emotion type different from the first emotion type.

[0194] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to obtain information about the environment in which the electronic device is located using the audio input. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the emotion information based further on the information about the environment.

[0195] According to one embodiment, the information regarding the tone of the voice signal may be first information regarding the tone of the voice signal. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to obtain second information regarding the tone of the voice signal by performing correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the emotion information based on the second information regarding the tone of the voice signal and the user information.

[0196] According to one embodiment, the instructions may cause the electronic device (101) to obtain information about the environment in which it is located using the audio input when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to provide the response to the voice signal based on providing the information about the environment and the information about the voice signal to the at least one target artificial intelligence model when executed individually or collectively by the at least one processor.

[0197] According to one embodiment, the electronic device may further include a communication circuit. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to receive sensing information from a wearable device connected to the electronic device through the communication circuit. The instructions, when executed individually or collectively by the at least one processor, may further cause the electronic device to identify emotion information based on the sensing information.

[0198] According to one embodiment, the instructions may cause the electronic device to generate a prompt including list information for the candidate artificial intelligence models and sentiment information when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine the at least one target artificial intelligence model among the candidate artificial intelligence models by providing the prompt to the target decision model when executed individually or collectively by the at least one processor.

[0199] According to one embodiment, a first portion of the candidate artificial intelligence models may be included within the electronic device. A second portion of the candidate artificial intelligence models may be included within at least one server device different from the electronic device.

[0200] According to one embodiment, the electronic device may further include a communication circuit. The instructions may cause the electronic device to identify, among at least one external electronic device registered with the electronic device, an external electronic device for executing a function according to the voice signal, based on providing information regarding the voice signal to the at least one target artificial intelligence model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to transmit a control signal for executing the function to the external electronic device through the communication circuit when executed individually or collectively by the at least one processor.

[0201] According to one embodiment, the user information may include previous emotion information that was identified before the emotion information was identified.

[0202] According to one embodiment, the user information may include at least one of conversation data related to a message application within the electronic device (101) collected within the time interval defined according to the reception of the audio input, or movement data of the electronic device (101) collected within the time interval defined according to the reception of the audio input.

[0203] A method performed in an electronic device having a microphone as described above may include an operation of receiving an audio input through the microphone. The method may include an operation of obtaining information regarding the tone of a voice signal included in the audio input and user information collected within a time interval defined by the reception of the audio input. The method may include an operation of identifying emotion information based on the information regarding the tone of the voice signal and the user information. The method may include an operation of determining at least one target artificial intelligence model among a plurality of candidate artificial intelligence models based on the emotion information. The method may include an operation of providing a response to the voice signal based on providing information regarding the voice signal to the at least one target artificial intelligence model.

[0204] According to one embodiment, the at least one target AI model may include at least one generative AI model. The method may include an operation of obtaining a prompt indicating at least one function by providing the information regarding the voice signal to the at least one generative AI model. The method may include an operation of obtaining response information using an execution result according to the execution of the at least one function. The method may include an operation of providing the response by processing the response information in a manner according to the emotion information.

[0205] According to one embodiment, the at least one target AI model may include at least one generative AI model. The method may include the operation of obtaining a prompt including an output method according to the emotion information by providing the information regarding the voice signal to the at least one generative AI model. The method may include the operation of providing the response having a data size according to the output method using the prompt.

[0206] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. The information regarding the voice signal may include text representing the voice signal and may include a prompt generated in a manner according to the at least one generative artificial intelligence model.

[0207] According to one embodiment, the electronic device may further include a speaker. The method may include an operation of outputting the response set to a first tone through the speaker based on identifying the emotion information corresponding to a first emotion type. The method may include an operation of outputting the response set to a second tone different from the first tone through the speaker based on identifying the emotion information corresponding to a second emotion type different from the first emotion type.

[0208] According to one embodiment, the electronic device may further include a display. The method may include an operation of displaying the response as an execution screen through the display based on identifying the emotion information corresponding to a first emotion type. The method may include an operation of displaying the response as a UI (user interface) object through the display based on identifying the emotion information corresponding to a second emotion type different from the first emotion type.

[0209] According to one embodiment, the method may include an operation of obtaining information about the environment in which the electronic device is located using the audio input. The method may further include an operation of identifying the emotion information based on the information about the environment.

[0210] According to one embodiment, the information regarding the tone of the voice signal may be first information regarding the tone of the voice signal. The method may include an operation of obtaining second information regarding the tone of the voice signal by performing correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal. The method may include an operation of identifying the emotion information based on the second information regarding the tone of the voice signal and the user information.

[0211] According to one embodiment, the method may include an operation of obtaining information about the environment in which the electronic device (101) is located using the audio input. The method may include an operation of providing the response to the voice signal based on providing the information about the environment and the information about the voice signal to the at least one target artificial intelligence model.

[0212] According to one embodiment, the electronic device may further include a communication circuit. The method may include the operation of receiving sensing information from a wearable device connected to the electronic device through the communication circuit. The method may further include the operation of identifying emotion information based on the sensing information.

[0213] According to one embodiment, the method may include the operation of generating a prompt including list information for the candidate artificial intelligence models and the sentiment information. The method may include the operation of determining at least one target artificial intelligence model among the candidate artificial intelligence models by providing the prompt to a target determination model.

[0214] According to one embodiment, a first portion of the candidate artificial intelligence models may be included within the electronic device. A second portion of the candidate artificial intelligence models may be included within at least one server device different from the electronic device.

[0215] According to one embodiment, the electronic device may further include a communication circuit. The method may include an operation of identifying, among at least one external electronic device registered with the electronic device, an external electronic device for executing a function according to the voice signal, based on providing information regarding the voice signal to at least one target artificial intelligence model. The method may include an operation of transmitting a control signal for executing the function to the external electronic device through the communication circuit.

[0216] According to one embodiment, the user information may include previous emotion information that was identified before the emotion information was identified.

[0217] According to one embodiment, the user information may include at least one of conversation data related to a message application within the electronic device (101) collected within the time interval defined according to the reception of the audio input, or movement data of the electronic device (101) collected within the time interval defined according to the reception of the audio input.

[0218] In a computer-readable storage medium storing one or more programs as described above, the one or more programs may include instructions that cause the electronic device to receive an audio input through the microphone when executed by the electronic device having a microphone. The one or more programs may include instructions that cause the electronic device to obtain information regarding the tone of a voice signal included in the audio input and user information collected within a time interval defined by the reception of the audio input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify emotion information based on the information regarding the tone of the voice signal and the user information when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine at least one target artificial intelligence model among a plurality of candidate artificial intelligence models based on the emotion information when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to provide a response to the voice signal based on providing information about the voice signal to the at least one target artificial intelligence model when executed by the electronic device.

[0219] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. The one or more programs may include instructions that cause the electronic device to obtain a prompt indicating at least one function by providing the information regarding the voice signal to the at least one generative artificial intelligence model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain response information by using the execution result of the execution of the at least one function when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the response by processing the response information in a manner according to the emotion information when executed by the electronic device.

[0220] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. The one or more programs may include instructions that cause the electronic device to obtain a prompt including an output method according to the emotion information by providing the information regarding the voice signal to the at least one generative artificial intelligence model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the response having a data size according to the output method by using the prompt when executed by the electronic device.

[0221] According to one embodiment, the at least one target artificial intelligence model may include at least one generative artificial intelligence model. The information regarding the voice signal may include text representing the voice signal and may include a prompt generated in a manner according to the at least one generative artificial intelligence model.

[0222] According to one embodiment, the electronic device may further include a speaker. The one or more programs may include instructions that cause the electronic device to output the response set to a first tone through the speaker, based on identifying the emotion information corresponding to a first emotion type when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to output the response set to a second tone different from the first tone through the speaker, based on identifying the emotion information corresponding to a second emotion type different from the first emotion type when executed by the electronic device.

[0223] According to one embodiment, the electronic device may further include a display. The one or more programs may include instructions that cause the electronic device to display the response as an execution screen through the display, based on identifying the emotion information corresponding to a first emotion type when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the response as a UI (user interface) object through the display, based on identifying the emotion information corresponding to a second emotion type different from the first emotion type when executed by the electronic device.

[0224] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain information about the environment in which the electronic device is located using the audio input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify the emotion information based further on the information about the environment when executed by the electronic device.

[0225] According to one embodiment, the information regarding the tone of the voice signal may be first information regarding the tone of the voice signal. The one or more programs may include instructions that cause the electronic device to obtain second information regarding the tone of the voice signal by performing correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify the emotion information based on the second information regarding the tone of the voice signal and the user information when executed by the electronic device.

[0226] According to one embodiment, the one or more programs may include instructions that cause the electronic device to obtain information about the environment in which the electronic device (101) is located using the audio input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the response to the voice signal based on providing the information about the environment and the information about the voice signal to the at least one target artificial intelligence model when executed by the electronic device.

[0227] According to one embodiment, the electronic device may further include a communication circuit. The one or more programs may include instructions that cause the electronic device to receive sensing information from a wearable device connected to the electronic device through the communication circuit when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify emotion information based further on the sensing information when executed by the electronic device.

[0228] According to one embodiment, the one or more programs may include instructions that cause the electronic device to generate a prompt including list information for the candidate artificial intelligence models and the sentiment information when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine the at least one target artificial intelligence model among the candidate artificial intelligence models by providing the prompt to a target determination model when executed by the electronic device.

[0229] According to one embodiment, a first portion of the candidate artificial intelligence models may be included within the electronic device. A second portion of the candidate artificial intelligence models may be included within at least one server device different from the electronic device.

[0230] According to one embodiment, the electronic device may further include a communication circuit. The one or more programs may include instructions that cause the electronic device to identify, among at least one external electronic device registered with the electronic device, an external electronic device for executing a function according to the voice signal, based on providing information regarding the voice signal to the at least one target artificial intelligence model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to transmit a control signal for executing the function to the external electronic device through the communication circuit when executed by the electronic device.

[0231] According to one embodiment, the user information may include previous emotion information that was identified before the emotion information was identified.

[0232] According to one embodiment, the user information may include at least one of conversation data related to a message application within the electronic device (101) collected within the time interval defined according to the reception of the audio input, or movement data of the electronic device (101) collected within the time interval defined according to the reception of the audio input.

[0233] For one or more embodiments, at least one of the components described in one or more of the prior art drawings may be configured to perform one or more operations, techniques, processes and / or methods as described in the present disclosure. For example, a processor (e.g., a baseband processor) described in the present disclosure in relation to one or more of the prior art drawings may be configured to operate according to one or more examples described in the present disclosure. As another example, circuits associated with user equipment (UE), a base station, a network element, etc., as described above in relation to one or more of the prior art drawings may be configured to operate according to one or more examples described herein.

[0234] Any of the embodiments described above may be combined with any other embodiment (or combination of embodiments) unless otherwise explicitly stated. The foregoing description of one or more embodiments is for illustrative and explanatory purposes only, and is not intended to limit or exhaust the scope of the embodiments in the exact form disclosed. Modifications and variations are possible in light of the foregoing teachings or may be obtained from the practice of various embodiments.

[0235] The electronic devices according to the various embodiments disclosed in this document may be of various forms. The electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, electronic devices, or consumer electronics. The electronic devices according to the embodiments of this document are not limited to the devices described above.

[0236] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0237] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0238] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101) of FIG. 1). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0239] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0240] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, microphone; Memory comprising one or more storage media for storing instructions; and It includes at least one processor comprising a processing circuit, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Receive audio input through the above microphone, and Information regarding the tone of a voice signal included in the above audio input and user information collected within a time interval defined according to the reception of the above audio input are obtained, Based on the information regarding the tone of the voice signal and the user information, emotional information is identified, and Based on the above-mentioned sentiment information, at least one target AI model is determined among a plurality of candidate AI models, and Causing to provide a response to the voice signal based on providing information about the voice signal to the at least one target artificial intelligence model. Electronic device.

2. In Claim 1, The above-mentioned at least one target AI model includes at least one generative AI model, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: By providing the information regarding the voice signal to the at least one generative artificial intelligence model, a prompt indicating at least one function is obtained, and Using the execution result of the execution of at least one of the above functions, response information is obtained, and By processing the above response information in a manner according to the above emotion information, causing to provide the above response, Electronic device.

3. In Claim 1, The above-mentioned at least one target AI model includes at least one generative AI model, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: By providing the information regarding the voice signal to the at least one generative artificial intelligence model, a prompt including an output method according to the emotion information is obtained, and Causing to provide the response having a data size according to the output method using the above prompt, Electronic device.

4. In Claim 1, The above-mentioned at least one target AI model includes at least one generative AI model, and The information regarding the voice signal includes text representing the voice signal and includes a prompt generated in a manner according to the at least one generative artificial intelligence model. Electronic device.

5. In Claim 1, Includes more speakers, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on identifying the emotional information corresponding to the first emotional type, the response set to the first tone is output through the speaker, and Based on identifying the emotion information corresponding to a second emotion type different from the first emotion type, causing the response set to a second tone different from the first tone to be output through the speaker. Electronic device.

6. In Claim 1, Includes more displays, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on identifying the emotion information corresponding to the first emotion type, the response is displayed as an execution screen through the display, and Based on identifying the emotion information corresponding to a second emotion type different from the first emotion type, the response is to be displayed as a UI (user interface) object through the display, and Electronic device.

7. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Using the above audio input, information about the environment in which the electronic device is located is obtained, and Causing to identify the emotional information based further on the information regarding the environment above, Electronic device.

8. In Claim 7, The information regarding the tone of the voice signal is the first information regarding the tone of the voice signal, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: By performing correction processing according to the information regarding the environment on the first information regarding the tone of the voice signal, second information regarding the tone of the voice signal is obtained, and Causing to identify the emotion information based on the second information regarding the tone of the voice signal and the user information. Electronic device.

9. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Using the above audio input, information about the environment in which the electronic device is located is obtained, and Based on providing the above information regarding the environment and the above information regarding the voice signal to the at least one target artificial intelligence model, causing to provide the above response to the voice signal, Electronic device.

10. In Claim 1, It further includes a communication circuit, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Receiving sensing information from a wearable device connected to the electronic device through the communication circuit, and Causing to identify the emotion information based further on the above sensing information, Electronic device.

11. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Generate a prompt including list information for the above candidate AI models and the above sentiment information, and By providing the above prompt to the target decision model, causing it to determine the at least one target AI model among the candidate AI models, Electronic device.

12. In Claim 1, A first part of the above candidate artificial intelligence models is included in the electronic device, and A second part of the above candidate artificial intelligence models is included in at least one server device other than the electronic device, Electronic device.

13. In Claim 1, It further includes a communication circuit, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on providing the information regarding the voice signal to the at least one target artificial intelligence model, an external electronic device for executing a function according to the voice signal is identified among at least one external electronic device registered with the electronic device, and Causing the external electronic device to transmit a control signal for the execution of the above function through the communication circuit, Electronic device.

14. In Claim 1, The above user information includes previous emotion information identified before the above emotion information was identified, Electronic device.

15. In Claim 1, The above user information includes at least one of conversation data related to a message application within the electronic device collected within the time interval defined according to the reception of the audio input, or motion data of the electronic device collected within the time interval defined according to the reception of the audio input. Electronic device.