Electronic device responding to user utterance, operating method thereof, and recording medium
The integration of gaze detection with voice input in AI devices optimizes conversational interactions by ensuring responses are delivered only when the user is visually engaged, enhancing user experience and interaction efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-07-29
- Publication Date
- 2026-04-30
AI Technical Summary
Existing AI devices lack the ability to effectively integrate user gaze detection with voice input to enhance conversational interactions, leading to suboptimal user experience and inefficient response management.
An electronic device equipped with a camera to detect user gaze direction in conjunction with voice input, allowing conditional responses based on gaze confirmation, thereby enhancing interaction efficiency.
Enables intelligent and context-aware conversational responses by ensuring responses are delivered only when the user is visually engaged with the device screen, improving user interaction and reducing unnecessary interactions.
Smart Images

Figure KR2025011244_30042026_PF_FP_ABST
Abstract
Description
Electronic device responding to user speech, method of operation thereof, and recording medium
[0001] The present disclosure relates to an electronic device that responds to user speech according to one embodiment, a method of operation thereof, and a recording medium.
[0002] With the rapid advancement of AI (artificial intelligence) technology, various AI agent applications and services are being developed, and the related market is growing explosively. Beyond massive language models, AI is expected to expand infinitely in the scope of its application across diverse aspects of life. This includes not only providing search results but also creating new content by combining existing data or acting as a substitute for actual experts by being trained with specialized knowledge in specific fields. Users can obtain what they want through conversation using AI agents. Consequently, this drastically reduces the time required to install applications and search for desired content across various applications or websites, in accordance with existing UX.
[0003] The above information may be provided as background information to aid in understanding the present disclosure. No claim or determination has been made as to which of the above constitutes prior art related to the present disclosure.
[0004] The aspects of the present disclosure are intended to solve at least the problems and / or disadvantages mentioned above and to provide at least the advantages described below. Accordingly, one aspect of the present disclosure is to provide an electronic device that responds to user utterance, a method of operating the same, and a recording medium.
[0005] Additional aspects will be presented in part in the following description, and in part may become apparent from the description or be learned by practicing the presented embodiments.
[0006] According to one embodiment, the electronic device may include a housing forming the exterior of the electronic device, a display disposed on a first surface of the housing, a camera disposed in the direction facing the display, at least one processor including a microphone and a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by at least one processor, the electronic device may cause the electronic device to operate in a conversational mode in which a voice agent executed by the electronic device interacts with a voice input received through the microphone. When the instructions are executed individually or collectively by at least one processor, the electronic device may cause the electronic device to check whether at least one condition for outputting a first response associated with the first voice input is satisfied after receiving the first voice input. When the instructions are executed individually or collectively by at least one processor, the electronic device may cause the electronic device to output the first response associated with the first voice input based on the satisfaction of the at least one condition for outputting the first response. When the above instructions are executed individually or collectively by at least one processor, the electronic device may be caused to receive a second voice input following the first voice input instead of outputting the first response, based on the fact that the at least one condition for outputting the first response is not satisfied. The at least one condition for outputting the first response may include an operation of confirming that the user's gaze of the electronic device, acquired using the camera, moves toward the screen of the display.
[0007] According to one embodiment, a method of operating an electronic device may include an operation in which a voice agent executed by the electronic device interacts with a voice input received through a microphone, operating in a conversational mode. The method may include an operation of checking whether at least one condition for outputting a first response related to the first voice input is satisfied after receiving a first voice input. The method may include an operation of outputting the first response related to the first voice input based on the satisfaction of the at least one condition for outputting the first response. The method may include an operation of receiving a second voice input following the first voice input instead of outputting the first response based on the failure to satisfy the at least one condition for outputting the first response. The at least one condition for outputting the first response may include an operation of checking that the gaze of the user of the electronic device, acquired using a camera, moves toward the screen of the display.
[0008] According to one embodiment, in a non-transitory computer-readable recording medium for storing instructions, the instructions may cause the electronic device to perform at least one operation when executed individually or collectively by at least one processor of the electronic device. The at least one operation may include an operation of operating in a conversational mode in which a voice agent executed by the electronic device interacts with a voice input received through a microphone. The at least one operation may include an operation of checking whether at least one condition for outputting a first response associated with the first voice input is satisfied after receiving a first voice input. The at least one operation may include an operation of outputting the first response associated with the first voice input based on the satisfaction of the at least one condition for outputting the first response. The at least one operation may include an operation of receiving a second voice input following the first voice input instead of outputting the first response based on the failure of the at least one condition for outputting the first response to be satisfied. The at least one condition for outputting the first response may include an operation to confirm that the user's gaze of the electronic device, acquired using a camera, moves to the screen of the display.
[0009] Other aspects, advantages, and key features of the present disclosure will become apparent to those skilled in the art from the following detailed description, together with reference to the accompanying drawings.
[0010] Other aspects, features, and advantages of the above and specific embodiments of the present disclosure will become more apparent from the following description, which is referenced together with the accompanying drawings.
[0011] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment of the present disclosure.
[0012] FIG. 2a is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0013] FIG. 2b is a drawing illustrating an electronic device according to one embodiment of the present disclosure.
[0014] FIG. 2c is a drawing for illustrating a generative artificial intelligence system according to one embodiment of the present disclosure.
[0015] FIG. 3 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0016] FIG. 4a is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0017] FIG. 4b is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0018] FIG. 4c is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0019] FIG. 4d is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0020] FIG. 4e is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0021] FIG. 4f is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0022] FIG. 5 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0023] FIG. 6 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0024] FIG. 7 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0025] FIG. 8 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0026] FIG. 9 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0027] FIG. 10 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0028] FIG. 11 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0029] FIG. 12 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0030] FIG. 13 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0031] FIG. 14 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0032] FIG. 15 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0033] FIG. 16 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0034] FIG. 17 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0035] FIG. 18 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0036] FIG. 19 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0037] FIG. 20 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0038] FIG. 21 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0039] FIG. 22 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0040] FIG. 23 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0041] FIG. 24 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0042] FIG. 25 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0043] FIG. 26 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0044] FIG. 27 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0045] FIG. 28 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0046] FIG. 29 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0047] FIG. 30 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0048] FIG. 31 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0049] FIG. 32 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0050] FIG. 33 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0051] The same reference number is used to represent the same elements in the drawings.
[0052] With reference to the accompanying drawings, the following description is provided to facilitate a comprehensive understanding of various embodiments of the present invention as defined by the claims and their equivalents. While this description includes various specific details to aid such understanding, they should be considered merely illustrative. Accordingly, those skilled in the art will recognize that various changes and modifications to the various embodiments described herein may be made without departing from the scope and spirit of the present invention. Furthermore, for clarity and brevity, descriptions of well-known functions and configurations may be omitted.
[0053] The terms and words used in the following description and claims are not limited to their bibliographic meanings but are used by the inventor merely to facilitate a clear and consistent understanding of the invention. Accordingly, it will be apparent to those skilled in the art that the following description of various embodiments of the invention is provided for illustrative purposes only and is not intended to limit the scope of the invention as defined by the appended claims and their equivalents.
[0054] The singular forms "one," "one," and "above" should be understood to include the plural form unless the context clearly indicates otherwise. Accordingly, for example, the expression "component surface" includes the fact that such surface may be one or more.
[0055] It should be understood that the blocks of each flowchart and combinations of flowcharts can be executed by one or more computer programs containing instructions. One or more computer programs as a whole may be stored in a single memory device, or one or more computer programs may be divided into multiple parts and stored in multiple different memory devices.
[0056] All functions or tasks described herein may be processed by a single processor or a combination of processors. A single processor or a combination of processors is a circuit that performs processing and includes an application processor (AP, e.g., central processing unit (CPU)), a communication processor (CP, e.g., a modem), a graphics processing unit (GPU), a neural processing unit (NPU) (e.g., an artificial intelligence (AI) chip), a wireless fidelity (Wi-Fi) chip, and Bluetooth. ® (Bluetooth ® It includes circuits such as ) chips, Global Positioning System (GPS) chips, Near Field Communication (NFC) chips, connectivity chips, sensor controllers, touch controllers, fingerprint sensor controllers, display driver integrated circuits (ICs), audio codec chips, Universal Serial Bus (USB) controllers, camera controllers, image processing ICs, microprocessor units (MPUs), system-on-chip (SoCs), ICs, etc.
[0057] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to one embodiment of the present disclosure.
[0058] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0059] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0060] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0061] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0062] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0063] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0064] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0065] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0066] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0067] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor. In one embodiment, the sensor module (176) may include sensor circuitry. In one embodiment, the sensor module (176) may include a first sensor, a second sensor, and / or a third sensor. In one embodiment, the sensor circuitry may include a first sensor, a second sensor, and / or a third sensor.
[0068] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0069] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0070] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0071] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0072] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0073] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0074] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G (fifth-generation) network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0075] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G (fourth-generation) networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0076] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0077] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0078] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0079] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0080] The operations of the electronic device (101) can be described in detail with reference to the aforementioned embodiments (e.g., embodiments of FIG. 1) and the embodiments described below (e.g., embodiments of FIG. 2a to 2c, 3, 4a to 4f, and 5 to 33). Each embodiment is disclosed in a separate drawing and a separate paragraph, but this is for convenience of explanation only, and at least some of the aforementioned embodiments and at least some of the embodiments described below may be applied together. At least some of the aforementioned embodiments and at least some of the embodiments described below may be omitted.
[0081] In this document, the electronic device (101) performing a specific operation may mean that a processor (120), such as a microcontrolling unit (MCU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a microprocessor, or an application processor (AP), performs a specific operation. According to one embodiment, the processor (120) may include a processing circuit. The electronic device (101) performing a specific operation may mean that the processor (120) controls other hardware to perform a specific operation. The electronic device (101) performing a specific operation may mean that the processor (120) or other hardware is caused to perform a specific operation as at least one instruction for performing a specific operation, which was stored in the storage circuit (e.g., memory (130)) of the electronic device (101), is executed. The at least one instruction stored in the memory (130) of the electronic device (101) may cause the electronic device (101) to perform at least one operation, either individually or collectively, when executed by the processor (120).
[0082] Functions related to artificial intelligence according to the present disclosure may be operated through a processor (120) and a memory (130). The processor (120) may be composed of one or more processors (120). In this case, the one or more processors (120) may be general-purpose processors such as a CPU, AP, DSP (digital signal processor), etc., graphics-dedicated processors such as a GPU, VPU (vision processing unit), or artificial intelligence-dedicated processors such as an NPU. The one or more processors (120) control input data to be processed according to predefined operation rules or artificial intelligence models stored in the memory (130). Alternatively, if the one or more processors (120) are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0083] A predefined rule of action or an artificial intelligence model may be said to be created through learning. Here, being created through learning means that a predefined rule of action or an artificial intelligence model is created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm to perform a desired characteristic (or purpose). The artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of operations of the previous layer and the multiple weights. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is performed, or it may be performed through a separate server (e.g., server (108) of FIG. 1) and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0084] Multiple weights possessed by multiple neural network layers of an artificial intelligence model can be optimized based on the learning results of the AI model. For example, multiple weights can be updated so that the loss or cost value obtained from the AI model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0085] In the method of operation of an electronic device (101) according to the present disclosure, the electronic device (101) can recognize voice. According to one embodiment, the electronic device (101) can recognize a user's voice by receiving an analog voice signal through an input module (150) (e.g., a microphone) and can interpret the intent of the utterance included in the recognized voice. According to one embodiment, the electronic device (101) can recognize audio included in a video and interpret the intent of the voice included in the recognized audio. The electronic device (101) can convert a voice portion into computer-readable text using an automatic speech recognition (ASR) model. It can obtain the intent of the utterance by interpreting the converted text using a natural language understanding (NLU) model. Here, the ASR model or the NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by a processor (120) (e.g., an artificial intelligence dedicated processor) designed with a hardware structure specialized for processing the artificial intelligence model. The artificial intelligence model may be created through learning. Linguistic understanding is a technology that recognizes, applies, and processes human language and text, and includes natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis.
[0086] In the method of operation of an electronic device (101) according to the present disclosure, the electronic device (101) can recognize an image. According to one embodiment, the electronic device (101) can obtain output data that recognizes an image by using image data as input data for an artificial intelligence model. The artificial intelligence model can be created through learning. Visual understanding is a technology that recognizes and processes objects like human vision, and includes object recognition, object tracking, image retrieval, human recognition, scene recognition, spatial understanding (3D reconstruction / localization), image enhancement, etc.
[0087] FIG. 2a is a block diagram of an electronic device according to one embodiment of the present disclosure. FIG. 2a may be described based on the embodiments of FIG. 1, the embodiments of FIG. 2b, the embodiments of FIG. 2c, the embodiments of FIG. 3, and the embodiments described below. FIG. 2b is a diagram illustrating an electronic device according to one embodiment of the present disclosure. FIG. 2c is a diagram for illustrating a generative artificial intelligence system according to one embodiment of the present disclosure. FIG. 3 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0088] Referring to FIG. 2a, as described above, the electronic device (101) may include a processor (120) and a memory (130).
[0089] Referring to FIG. 2b, according to one embodiment, an electronic device (101) may include a housing (200) that forms the exterior of the electronic device (101). For example, components of the electronic device (101) (e.g., components of FIG. 1 and components of FIG. 2a) may be placed in the housing (200).
[0090] Referring to FIG. 2a, according to one embodiment, an electronic device (101) may include a microphone (210) (e.g., the microphone of the input module (150) of FIG. 1). The microphone (210) may be configured to acquire voice data. The electronic device (101) may acquire voice data through the microphone (210). The voice data may include data for the voice of a single speaker or data for the voices of multiple speakers, and there are no limitations on the type of voice data or the type of information included in the voice data. There are no limitations on the number and location of the microphone (210). According to one embodiment, the electronic device (101) may measure the distance between the electronic device (101) and the speaker using a plurality of microphones (210). For example, the electronic device (101) may measure the distance between the electronic device (101) and the speaker based on the difference in the time at which the plurality of microphones (210) receive the voice. According to one embodiment, the electronic device (101) may measure the distance between the electronic device (101) and the speaker using a sensor (e.g., a second sensor (232)) as described below. According to one embodiment, the electronic device (101) may determine the direction in which the speaker is located from the electronic device (101) using a plurality of microphones (210). For example, the electronic device (101) may determine the direction in which the speaker is located from the electronic device (101) based on the difference in the time when the plurality of microphones (210) receive the voice.
[0091] Referring to FIG. 2a, according to one embodiment, an electronic device (101) may include a display (250) (e.g., the display of the display module (160) of FIG. 1). The display (250) may be configured to display a screen. The electronic device (101) may control the display (250) to display a screen. The display (250) may be configured to receive input (e.g., touch input). The electronic device (101) may receive input (e.g., touch input) through the display (250). According to one embodiment, referring to FIG. 2b, the display (250) (e.g., the display of the display module (160) of FIG. 1) may be placed on a first surface (e.g., the front) of the housing (200). For example, the electronic device (101) may include a display (250) (e.g., the display of the display module (160) of FIG. 1) placed on a first surface (e.g., the front) of the housing (200). For example, the direction in which the display (250) (e.g., the display of the display module (160) of FIG. 1) faces may be defined as the front of the electronic device (101). As the user looks at the front of the electronic device (101), the user may gaze at the display (250) (e.g., the display of the display module (160) of FIG. 1) placed on the first surface (e.g., the front) of the housing (200).
[0092] Referring to FIG. 2a, according to one embodiment, the electronic device (101) may include a camera (220) (e.g., the camera of the camera module (180) of FIG. 1). The camera (220) may be configured to acquire image data. There is no limit to the number of cameras (220). The electronic device (101) may acquire image data through the camera (220). The image data may include data of a single image frame or continuous data for a plurality of image frames. The image data may include information about the user's eyes, mouth, face, gestures, and / or facial expressions, and there is no limit to the type of image data and the type of information included in the image data. According to one embodiment, the camera (220) (e.g., a front camera) may be configured to acquire an image in the direction facing the display (250) (e.g., the display of the display module (160) of FIG. 1). For example, a camera (220) (e.g., a front camera) may be positioned in the direction in which the display (250) (e.g., the display of the display module (160) of FIG. 1) is facing. For example, the camera (220) (e.g., a front camera) may be configured to acquire an image in the direction in which the front of the electronic device (101) is facing. For example, referring to FIG. 2b, the electronic device (101) may acquire an image in the direction in which the display (250) (e.g., the display of the display module (160) of FIG. 1) is facing through the camera (220) (e.g., a front camera).
[0093] Referring to FIG. 2a, according to one embodiment, the electronic device (101) may include a rear camera (222) (e.g., the camera of the camera module (180) of FIG. 1). For example, the rear camera (222) may be configured to acquire an image in the opposite direction to the direction in which the display (250) (e.g., the display of the display module (160) of FIG. 1) is facing. For example, the rear camera (222) may be positioned in the opposite direction to the direction in which the display (250) (e.g., the display of the display module (160) of FIG. 1) is facing. The rear camera (222) may be configured to acquire image data. For example, the rear camera (222) may be configured to acquire an image in the direction in which the rear of the electronic device (101) is facing. There is no limit to the number of rear cameras (222). The electronic device (101) may acquire image data through the rear camera (222).
[0094] According to one embodiment, the on (e.g., enabled) and off (e.g., disabled) of the rear camera (222) may be independent of the conversation mode described below. The conversation mode may be a mode in which a voice agent (e.g., an artificial intelligence model) executed by the electronic device (101) interacts with voice input received through the microphone (210). The conversation mode will be described below. For example, the on (e.g., enabled) and off (e.g., disabled) of the camera (220) (e.g., front camera) may be determined by the operation of the conversation mode, but the on and off of the rear camera (222) may be independent of the conversation mode. For example, based on the activation of the conversation mode, the electronic device (101) may enable the camera (220) (e.g., front camera). For example, the electronic device (101) may disable the camera (220) based on checking conditions for disabling the camera (220) while the conversation mode is being performed. For example, the electronic device (101) may enable the camera (220) based on checking specific conditions (e.g., checking the standing state or holding state described below) while the conversation mode is being performed with the camera (220) disabled. For example, the electronic device (101) may disable the camera (220) based on the conversation mode being disabled while the camera (220) is enabled.
[0095] Referring to FIG. 2a, according to one embodiment, an electronic device (101) may include a first sensor (231) (e.g., a sensor of the sensor module (176) of FIG. 1). The first sensor (231) may be configured to identify the posture (e.g., tilted direction, degree of tilt) and / or movement (e.g., movement along the X-axis, Y-axis, and Z-axis) of the electronic device (101). For example, the first sensor (231) may include a gyroscope, an accelerometer, and / or a 6-axis sensor, but there is no limitation on the type and number of the first sensor (231). The first sensor (231) may be configured to acquire sensing data to identify the posture and / or movement of the electronic device (101). The electronic device (101) may identify the posture and / or movement of the electronic device (101) using the sensing data (e.g., a sensing value) of the first sensor (231). There is no limit to the number of first sensors (231), and the electronic device (101) can identify the posture and / or movement of the electronic device (101) based on a combination of sensing data from a plurality of first sensors (231).
[0096] Referring to FIG. 2a, according to one embodiment, an electronic device (101) may include a second sensor (232) (e.g., a sensor of the sensor module (176) of FIG. 1). The second sensor (232) may be configured to measure the distance between the electronic device (101) and a user. A configuration for measuring the distance between the electronic device (101) and a user may be referred to as the second sensor (232). For example, the second sensor (232) may include a distance sensor (e.g., a ToF (time of flight) sensor), but there are no limitations on the type and number of the second sensor (232). The second sensor (232) may be configured to acquire sensing data for measuring the distance between the electronic device (101) and a user. The electronic device (101) may measure the distance between the electronic device (101) and a user by using the sensing data (e.g., a sensing value) of the second sensor (232). There is no limit to the number of second sensors (232), and the electronic device (101) can measure the distance between the electronic device (101) and the user based on a combination of sensing data from a plurality of second sensors (232). As described above, according to one embodiment, the electronic device (101) may measure the distance between the electronic device (101) and the user by using a microphone (210) (e.g., a plurality of microphones (210)). According to one embodiment, the electronic device (101) may measure the distance between the electronic device (101) and the user by using the sensing value of the second sensor (232) and voice data from the microphone (210).
[0097] Referring to FIG. 2a, according to one embodiment, the electronic device (101) may include a speaker (240) (e.g., the speaker of the sound output module (155) of FIG. 1). The speaker (240) may be configured to output sound. The electronic device (101) may output sound using the speaker. For example, the electronic device (101) may output a response (e.g., sound) through the speaker (240) in response to a user's speech.
[0098] Referring to FIG. 2a, according to one embodiment, an electronic device (101) may include a communication circuit (260) (e.g., a communication circuit of the communication module (190) of FIG. 1). The communication circuit (260) may be configured to perform communication with an external device (e.g., the electronic device (102) of FIG. 1, the electronic device (104), the server (108), the wearable device (1900) of FIG. 19, or the external device (3310, 3320, 3330) of FIG. 33). The electronic device (101) may perform communication with an external device (e.g., the electronic device (102) of FIG. 1, the electronic device (104), the server (108), the wearable device (1900) of FIG. 19, or the external device (3310, 3320, 3330) of FIG. 33) using the communication circuit (260).
[0099] Referring to FIG. 2c, a generative artificial intelligence system may be described according to one embodiment. The generative artificial intelligence system may include a user query / response interface (271), an AI framework (280), a knowledge repository (272), an application / service component (273), and / or a generative AI model (274). Referring to FIG. 2c, the user query / response interface (271) may receive user input. The user input may be in the form of natural language, images, and / or videos, but is not limited thereto. Additionally, context information may be transmitted along with the user input. The context information may include various additional information at the time of user input. For example, the additional information may include information about the application currently being used by the user or the user's location information. Additionally, the user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Additionally, user input may be in a non-natural language form, such as selecting a menu. The user query / response interface (271) can output results from the generative artificial intelligence system to the user. The output may be in a natural language form or a specific content form, and may also be provided in the form of actions requested by the user. The user query / response interface (271) can output results from the generative artificial intelligence system to the user. The output may be in a natural language form or a specific content form, and may also be provided in the form of actions requested by the user.
[0100] The AI framework (280) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query.
[0101] User input received from the user query / response interface (271) can be transmitted to a prompt design component (281). The prompt design component (281) can be used to generate prompts suitable for inputting user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (281) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (281) can generate prompts by accessing a knowledge component containing user preference data, a prompt library, and prompt examples based on user input, and can transmit the generated prompts to the LLM or LMM.
[0102] The API / Plug-in management component (282) can perform the role of communicating with external information when there is a request for additional information when user input is passed as input to a generative model. The API / Plug-in management component (282) establishes a channel to communicate with the outside of the AI Interface via the API, and through the established channel, it can enable access to various data sources (e.g., knowledge repository (272)). Additionally, the API / Plug-in management component (282) can request the application / service component (273) via the API to perform an action that ultimately executes the user input, rather than an intermediate result, in the application or service. The information obtained from the outside may be used to generate a prompt in the prompt design component (281) along with the user input, or it may be passed as input to the generative model.
[0103] The output modification component (or refiner component) (283) can fine-tune the output of the generative model. For example, the output modification component (283) can verify whether the content generated through the LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the output modification component (283) can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Furthermore, the output modification component (283) can configure and provide hints to the user to avoid unwanted output.
[0104] A generative AI model (274) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (274) may include models that generate images and / or models that generate language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples of models that use VAEs and Diffusion-based generative models with Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples of which include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs (large multimodal models) that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.
[0105] Referring to FIG. 3, an electronic device (101) that responds to user speech can be described according to one embodiment.
[0106] FIG. 3 illustrates a situation in which a first speaker (310) and a second speaker (320) are located around an electronic device (101). In FIG. 3 (a), the first speaker (310) may make an utterance requesting translation (e.g., "Translate it") to the electronic device (101). Upon receiving voice data for the utterance requesting translation (e.g., "Translate it"), the electronic device (101) may perform the translation requested by the first speaker (310). For example, in FIG. 3 (a), the electronic device (101) may output a response (e.g., "Yes") indicating that the user's request (e.g., translation request) has been recognized, and then output a response containing the translated sentence. In FIG. 3(b), the first speaker (310) may make a speech proposing to the second speaker (320) (e.g., "Minsu, shall we go have a highball later?"). The electronic device (101) may receive voice data regarding the speech proposing to the first speaker (310) (e.g., "Minsu, shall we go have a highball later?"). After the electronic device (101) receives voice data regarding the speech proposing to the first speaker (310) (e.g., "Minsu, shall we go have a highball later?") in FIG. 3(b), the situation of FIG. 3(c-1) or the situation of FIG. 3(c-2) may unfold. FIG. 3(c-1) may be a situation in which the electronic device (101) responds. In (c-1) of FIG. 3, the electronic device (101) may output a response (e.g., "Yes, did you call me?" or "I want to go drink a highball") upon receiving voice data for a speech suggested by the first speaker (310) (e.g., "Minsu, do you want to go drink a highball later?"). However, the response of the electronic device (101) in (c-1) of FIG. 3 may not be the response intended by the first speaker (310). FIG. 3 (c-1) may be a situation where the first speaker (310) wanted a response from the second speaker (320), but the electronic device (101) responded instead.In this case, the first speaker (310) and the second speaker (320) may become confused or not trust the electronic device (101). (c-2) of FIG. 3 may be a situation where the electronic device (101) does not respond. In (c-2) of FIG. 3, the electronic device (101) may not output a response to the voice data even though it has acquired voice data for the utterance suggested by the first speaker (310) (e.g., "Minsu, do you want to go drink a highball later?").
[0107] The following embodiments may describe embodiments in which the electronic device (101) does not operate as in (c-1) of FIG. 3, but operates as in (a) and (c-2) of FIG. 3.
[0108] The following drawings are separated into individual drawings for the convenience of explanation, and those skilled in the art will understand that at least some of the embodiments of the following drawings may be applied in conjunction with each other.
[0109] FIG. 4a is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 4a can be explained based on the embodiments of FIG. 1, 2a, 2b, 2c and FIG. 3, and the embodiments described below.
[0110] Referring to FIG. 4a, according to one embodiment, the electronic device (101) can determine the operation mode of the conversation mode (e.g., listening mode or response mode).
[0111] At least some of the operations of FIG. 4a may be omitted. The order of the operations of FIG. 4a may be changed. Operations other than those of FIG. 4a may be performed before, during, or after the operations of FIG. 4a.
[0112] Referring to FIG. 4a, in operation 401, according to one embodiment, the electronic device (101) may enter a conversation mode. The "conversation mode" may be a mode in which a user of the electronic device (101) and an artificial intelligence model (e.g., a voice agent) engage in a conversation. For example, the "conversation mode" may be a mode in which a voice agent executed by the electronic device (101) interacts with voice input (e.g., voice data) received through a microphone (210). The voice input may be data obtained through the microphone (210) based on the user's speech. For example, the operation modes of the conversation mode may include a listening mode and a response mode, the listening mode is described in operation 405 and the response mode is described in operation 407. While operating in the conversation mode, the electronic device (101) may engage in a conversation with the user using an artificial intelligence model (e.g., a voice agent). A conversation between an electronic device (101) and a user can be performed by the electronic device (101) acquiring voice data (e.g., voice input) regarding the user's speech according to the user's speech, verifying response information generated based on analysis of the voice data using an artificial intelligence model (e.g., voice agent), and outputting a response based on the response information. According to one embodiment, an on-device artificial intelligence model (e.g., voice agent) of the electronic device (101) may be used to perform at least some operations during the process of performing a conversation between the electronic device (101) and the user. According to one embodiment, an artificial intelligence model (e.g., voice agent) of the server (108) may be used to perform at least some operations during the process of performing a conversation between the electronic device (101) and the user.According to one embodiment, in order to perform at least some operations in the process of performing a conversation between an electronic device (101) and a user, an on-device artificial intelligence model (e.g., voice agent) of the electronic device (101) and an artificial intelligence model (e.g., voice agent) of the server (108) may be used individually or collectively. For example, an operation to generate response information based on an analysis of voice data may be performed by the on-device artificial intelligence model (e.g., voice agent) of the electronic device (101), or by the artificial intelligence model (e.g., voice agent) of the server (108), or individually or collectively by the on-device artificial intelligence model (e.g., voice agent) of the electronic device (101) and the artificial intelligence model (e.g., voice agent) of the server (108). Although examples have been given only for operations that generate response information, this is for convenience of explanation, and the on-device artificial intelligence model (e.g., voice agent) of the electronic device (101) and / or the artificial intelligence model (e.g., voice agent) of the server (108) may be used individually or collectively for other operations that require the intervention of an artificial intelligence model (e.g., voice agent).
[0113] According to one embodiment, in operation 401, the electronic device (101) may enter a conversation mode based on identifying an event that causes entry into a conversation mode with an artificial intelligence model (e.g., a voice agent). Examples of events that cause entry into a conversation mode are as follows. According to one embodiment, the electronic device (101) may enter a conversation mode based on a call to the electronic device (101). The call to the electronic device (101) may be a voice command that causes the electronic device (101) to enter a conversation mode. For example, the electronic device (101) may enter a conversation mode based on receiving voice data corresponding to an utterance including a wake word (e.g., "Hi Bixby"). According to one embodiment, the electronic device (101) may enter a conversation mode based on input to a button (e.g., a hardware button of the electronic device (101), or a software button displayed on the screen of the display (250) of the electronic device (101). According to one embodiment, the electronic device (101) may enter a conversation mode based on confirming that the electronic device (101) is mounted on a stand. For example, the electronic device (101) may confirm that the electronic device (101) is mounted on a stand based on sensing data (e.g., a sensing value) of at least one sensor (e.g., the first sensor (231) of FIG. 2a) configured to identify the posture and / or movement of the electronic device (101). Based on a setting that causes the electronic device (101) to enter a conversation mode when mounted on a stand, the electronic device (101) may enter a conversation mode based on confirming that the electronic device (101) is mounted on a stand.According to one embodiment, the electronic device (101) may enter a conversation mode based on receiving a request from an external device (e.g., the electronic device (102) of FIG. 1, the electronic device (104), the server (108), the wearable device (1900) of FIG. 19, or the external device (3310, 3320, 3330) of FIG. 33) through a communication circuit (260). For example, a user may input an input that causes the electronic device (101) to enter a conversation mode through a wearable device (e.g., the external device (3310, 3320, 3330) of FIG. 33) that is connected to the electronic device (101) through communication. Based on input (e.g., user input) through a wearable device (e.g., external device (3310, 3320, 3330) of FIG. 33) that is communication-connected to the electronic device (101), the electronic device (101) may receive a request from the wearable device (e.g., external device (3310, 3320, 3330) of FIG. 33) that causes entry into conversation mode. The events causing entry into conversation mode are exemplary, and there is no limitation on the types of events causing entry into conversation mode. The 401 operation may be understood as an operation to re-enter conversation mode after the conversation mode has ended or been stopped.
[0114] 403 In operation, according to one embodiment, the electronic device (101) may determine the operation mode of the conversation mode. The operation mode of the conversation mode may define the operation to be performed by the electronic device (101) during the conversation mode. As the operation mode of the conversation mode is determined, the operation performed by the electronic device (101) while the electronic device (101) is operating in the conversation mode may be determined. The operation mode of the conversation mode may include a listening mode and a response mode. As described below, a group mode and a basic mode may also be one of the operation modes of the conversation mode. For example, a combination of modes, such as a group mode and a response mode or a basic mode and a response mode, may also be an operation mode of the conversation mode. According to one embodiment, the electronic device (101) may determine the operation mode of the conversation mode based on conditions for determining the operation mode of the conversation mode. Specific conditions for determining the operation mode of the conversation mode will be described with reference to the embodiment of FIG. 5 and the embodiments of FIG. 6 through FIG. 33. In FIG. 4, operations performed in the listening mode or the response mode among the operation modes of the conversation mode can be described. While the electronic device (101) is operating in the conversation mode, the electronic device (101) can switch the operation mode of the conversation mode according to conditions. While the electronic device (101) is operating in the conversation mode, the electronic device (101) can switch to an operation mode corresponding to the changed condition (e.g., listening mode or response mode) based on confirming a change in the condition. For example, while the electronic device (101) is operating in the conversation mode, the electronic device (101) can switch from the listening mode to the response mode, or switch from the response mode to the listening mode. According to one embodiment, the electronic device (101) can determine the operation mode according to conditions when it first enters the conversation mode.For example, when the electronic device (101) first enters the conversation mode, it may check the conditions and determine an operating mode corresponding to the checked conditions. According to one embodiment, when the electronic device (101) first enters the conversation mode, it may determine an operating mode according to the settings. For example, if the listening mode (or response mode) is set as the default, the electronic device (101) may determine the listening mode (or response mode) as the operating mode according to the settings when it first enters the conversation mode. When the electronic device (101) first enters the conversation mode, after determining the listening mode (or response mode) as the operating mode according to the settings, it may check the conditions and switch to the response mode (or listening mode) according to the checked conditions.
[0115] In operation 405, according to one embodiment, the electronic device (101) may operate in a listening mode based on determining the listening mode as the operating mode. The operation of the listening mode may include an operation of acquiring voice data. While operating in the listening mode, the electronic device (101) may acquire voice data corresponding to a user's speech using a microphone (210). The operation of the listening mode may include an operation of storing the acquired voice data in a memory (130). While operating in the listening mode, the electronic device (101) may store the voice data acquired using the microphone (210) in the memory (130). The listening mode may be a mode that does not output a response. While operating in the listening mode, the electronic device (101) may not output a response. The voice data acquired (or stored) during the listening mode may be used in the response mode. The electronic device (101) may also use the voice data acquired (or stored) during the listening mode in the response mode. According to one embodiment, the electronic device (101) may not use voice data acquired (or stored) during the listening mode in the response mode.
[0116] In operation 407, according to one embodiment, the electronic device (101) may operate in a response mode based on determining the response mode as an operation mode. The operation of the response mode may include an operation of acquiring voice data. While operating in the response mode, the electronic device (101) may acquire voice data corresponding to a user's speech using a microphone (210). The operation of the response mode may include an operation of storing the acquired voice data in a memory (130). While operating in the response mode, the electronic device (101) may store the voice data acquired using the microphone (210) in the memory (130). The response mode may be a mode for outputting a response. While operating in the response mode, the electronic device (101) may output a response corresponding to the voice data. For example, the operation of the response mode may include an operation of verifying response information generated based on the voice data. Response information may be information generated based on the analysis of an artificial intelligence model (e.g., voice agent) on acquired voice data (e.g., an on-device artificial intelligence model of the electronic device (101) and / or an artificial intelligence model of the server (108). For example, the electronic device (101) may generate response information corresponding to voice data by analyzing voice data (e.g., voice data accumulated during conversation mode, or voice data acquired during response mode) using an on-device artificial intelligence model (e.g., voice agent) of the electronic device (101) during response mode. For example, the electronic device (101) may transmit voice data (e.g., voice data accumulated during conversation mode, or voice data acquired during response mode) to the server (108) and receive response information generated by the server (108) from the server (108). Voice data accumulated during conversation mode may be voice data acquired (or stored) while the electronic device (101) is operating in conversation mode, including listening mode and response mode.The electronic device (101) may use voice data accumulated during conversation mode or voice data acquired during response mode. The operation of the response mode may include an operation of outputting a response based on verified voice information. While operating in response mode, the electronic device (101) may output a response (e.g., sound) through the speaker (240) based on response information generated based on voice data.
[0117] Referring to the following embodiments, examples of operations determining the operation mode of the 403 operation of FIG. 4a will be described.
[0118] FIG. 4b is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 4b can be explained based on the embodiments of FIG. 1, 2a, 2b, 2c, and FIG. 3, the embodiments of FIG. 4a, and the embodiments described below.
[0119] At least some of the operations of FIG. 4b may be omitted. The order of the operations of FIG. 4b may be changed. Operations other than those of FIG. 4b may be performed before, during, or after the operations of FIG. 4b.
[0120] FIG. 4b may be a diagram for explaining the operations of FIG. 4a in a chronological order. For example, the operations of FIG. 4b may be embodiments of the operations of FIG. 4a. Any parts of the description of the operations of FIG. 4b that overlap with the description of the operations of FIG. 4a may be omitted.
[0121] Referring to FIG. 4b, in operation 421, according to one embodiment, the electronic device (101) may operate in a conversation mode. As previously described, the "conversation mode" may be a mode in which a voice agent executed by the electronic device (101) interacts with voice input received through the microphone (210). The voice input may be data obtained through the microphone (210) based on the user's speech. While operating in a conversation mode, the electronic device (101) may perform a conversation with the user using a voice agent (e.g., an artificial intelligence model). The conversation between the electronic device (101) and the user may be performed by the electronic device (101) acquiring voice data (e.g., voice input) for the user's speech according to the user's speech, verifying response information generated based on the analysis of the voice data (e.g., voice input) using an artificial intelligence model (e.g., a voice agent), and outputting a response based on the response information. The 421 operation can be understood by referring to the description of the 401 operation in FIG. 4a.
[0122] In operation 423, according to one embodiment, the electronic device (101) may receive voice input. For example, the electronic device (101) may receive voice input (e.g., voice data for a user's speech) through a microphone (210). The electronic device (101) may receive a first voice input, for example, at a first time. For example, the electronic device (101) may receive a first voice input through the microphone (210) based on a user's speech.
[0123] In operation 425, according to one embodiment, the electronic device (101) may check whether at least one condition for outputting a response is satisfied. "At least one condition for outputting a response" may be referred to as a "response condition." For example, "at least one condition for outputting a response" (e.g., "response condition") may include an operation of checking for the occurrence of an event that causes to output a response related to voice input. For example, the electronic device (101) may output a response related to voice input based on the satisfaction of at least one condition for outputting a response (e.g., response condition). For example, "at least one condition for outputting a response" (e.g., "response condition") may be a condition that causes to operate in the response mode of FIG. 4a. There is no limit to the number of at least one condition for outputting a response (e.g., response condition). According to one embodiment, the electronic device (101) may output a response related to a voice input based on at least one of at least one condition for outputting a response (e.g., a response condition) being verified. For example, after receiving a first voice input (e.g., a first voice input of a 423 operation), the electronic device (101) may verify whether at least one condition for outputting a first response related to the first voice input is satisfied.
[0124] In operation 427, according to one embodiment, the electronic device (101) may output a response related to a voice input based on at least one of at least one condition for outputting a response (e.g., response condition) being satisfied. For example, the electronic device (101) may perform operation 407 of FIG. 4a based on at least one of at least one condition for outputting a response (e.g., response condition) being satisfied. For example, the electronic device (101) may output a first response related to a first voice input based on at least one condition for outputting a first response being satisfied. For example, the electronic device (101) may check response information generated based on an analysis of a voice input (e.g., voice data) using an artificial intelligence model (e.g., voice agent) based on at least one condition for outputting a first response being satisfied, and output a response based on the response information. For example, at least one condition for outputting a first response may include an operation of confirming that the user's gaze of the electronic device (101), obtained using a camera (220), moves toward the screen of the display (250). For example, the electronic device (101) may confirm the user's gaze of the electronic device (101) using a camera (220), and output a first response related to a first voice input based on confirming that the user's gaze moves toward the screen of the display (250). There is no limitation on the type and number of at least one condition for outputting a response. In the embodiments of the drawings described below, at least one condition for outputting a response will be explained.
[0125] In operation 429, according to one embodiment, the electronic device (101) may receive a subsequent voice input (e.g., a second voice input following the first voice input) instead of outputting a response (e.g., a first response related to the first voice input) based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied. For example, the electronic device (101) may perform operation 405 of FIG. 4a based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied. For example, the electronic device (101) may not output a response (e.g., a first response related to the first voice input) based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied. For example, the electronic device (101) may receive a subsequent voice input (e.g., a second voice input following the first voice input) based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied. For example, the electronic device (101) may wait without outputting a response (e.g., a first response related to the first voice input) in order to receive a subsequent voice input (e.g., a second voice input following the first voice input) based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied.
[0126] According to one embodiment, the electronic device (101) may check whether at least one condition for outputting a second response related to the first voice input and the second voice input is satisfied after receiving a second voice input (e.g., the second voice input of the operation 429 of FIG. 4b following the first voice input) without outputting a first response related to the first voice input (e.g., the first voice input of the operation 423 of FIG. 4b). For example, the at least one condition for outputting the second response may be at least one condition for outputting the response of the operation 425 of FIG. 4b (e.g., a response condition). For example, the second response may be a response generated based on the second voice input. For example, the second response may be a response generated based on the first voice input and the second voice input. According to one embodiment, the electronic device (101) may output a second response related to a first voice input and a second voice input based on at least one condition (e.g., a response condition) for outputting a second response being satisfied. According to one embodiment, the electronic device (101) may output a second response related to a second voice input based on at least one condition (e.g., a response condition) for outputting a second response being satisfied. According to one embodiment, the electronic device (101) may receive a third voice input following the second voice input instead of outputting a second response based on at least one condition (e.g., a response condition) for outputting a second response not being satisfied. For example, the electronic device (101) may wait to receive a voice input following the second voice input instead of outputting a second response based on at least one condition (e.g., a response condition) for outputting a second response not being satisfied.
[0127] FIG. 4c is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 4c can be described based on the embodiments of FIG. 1, 2a, 2b, 2c and FIG. 3, the embodiments of FIG. 4a and FIG. 4b and the embodiments described below.
[0128] At least some of the operations of FIG. 4c may be omitted. The order of the operations of FIG. 4c may be changed. Operations other than those of FIG. 4c may be performed before, during, or after the operations of FIG. 4c.
[0129] FIG. 4c may be a diagram for explaining the operations of FIG. 4a in a chronological order. For example, the operations of FIG. 4c may be embodiments of the operations of FIG. 4a. Any parts of the description of the operations of FIG. 4c that overlap with the description of the operations of FIG. 4a may be omitted.
[0130] Referring to FIG. 4c, in operation 431, according to one embodiment, the electronic device (101) may operate in a conversational mode. Operation 431 can be understood by referring to the description of operation 421 in FIG. 4b and the description of operation 401 in FIG. 4a.
[0131] In operation 433, according to one embodiment, the electronic device (101) may receive voice input. Operation 433 can be understood by referring to the description of operation 423 of FIG. 4b. The electronic device (101) may receive a first voice input, for example, at a first time. For example, the electronic device (101) may receive the first voice input through a microphone (210) based on user speech.
[0132] In operation 435, according to one embodiment, the electronic device (101) may check whether a second condition (e.g., a listening condition) is satisfied. The second condition (e.g., a listening condition) may be at least one condition for receiving subsequent voice input instead of outputting a response. The second condition (e.g., a listening condition) may include an operation to check for the occurrence of an event that causes not to output a response related to the voice input. For example, the electronic device (101) may not output a response related to the voice input based on the second condition (e.g., a listening condition) being satisfied. For example, the electronic device (101) may receive subsequent voice input based on the second condition (e.g., a listening condition) being satisfied. For example, the second condition (e.g., a listening condition) may be a condition that causes to operate in the listening mode of FIG. 4a. There is no limit to the number of second conditions (e.g., a listening condition). According to one embodiment, the electronic device (101) may receive a subsequent voice input instead of outputting a response based on at least one of the second conditions (e.g., listening conditions) being confirmed. For example, the electronic device (101) may check whether the second condition (e.g., listening conditions) is satisfied after receiving the first voice input (e.g., the first voice input of a 433 operation).
[0133] In operation 437, according to one embodiment, the electronic device (101) may receive a subsequent voice input instead of outputting a response based on at least one of the second conditions (e.g., listening conditions) being satisfied. For example, the electronic device (101) may perform operation 405 of FIG. 4a based on at least one of the second conditions (e.g., listening conditions) being satisfied. For example, the electronic device (101) may not output a response based on the second condition (e.g., listening conditions) being satisfied. For example, the electronic device (101) may receive a subsequent voice input based on the second condition (e.g., listening conditions) being satisfied. For example, the second condition (e.g., listening conditions) may include an operation of confirming that the user's gaze, obtained using the camera (220), moves out of the screen of the display (250). For example, the electronic device (101) can use a camera (220) to check the user's gaze of the electronic device (101) and, based on checking that the user's gaze has moved out of the screen of the display (250), receive a subsequent voice input instead of outputting a response. There is no limit to the type and number of the second condition (e.g., listening condition). In the embodiments of the drawings described below, the second condition (e.g., listening condition) will be described.
[0134] 439 In operation, according to one embodiment, the electronic device (101) may check whether another condition (e.g., at least one condition for outputting the response of FIG. 4b (e.g., response condition)) is satisfied based on the fact that the second condition (e.g., listening condition) is not satisfied.
[0135] FIG. 4d is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 4d can be described based on the embodiments of FIG. 1, 2a, 2b, 2c and FIG. 3, the embodiments of FIG. 4a, 4b and 4c and the embodiments described below.
[0136] At least some of the operations of FIG. 4d may be omitted. The order of the operations of FIG. 4d may be changed. Operations other than those of FIG. 4d may be performed before, during, or after the operations of FIG. 4d.
[0137] FIG. 4d may be a drawing for explaining the operations of FIG. 4a in a chronological order. For example, the operations of FIG. 4d may be embodiments of the operations of FIG. 4a. Any parts of the description of the operations of FIG. 4d that overlap with the description of the operations of FIG. 4a may be omitted.
[0138] Referring to FIG. 4d, in operation 441, according to one embodiment, the electronic device (101) may operate in a conversational mode. Operation 441 can be understood by referring to the description of operation 421 in FIG. 4b and the description of operation 401 in FIG. 4a.
[0139] In operation 443, according to one embodiment, the electronic device (101) may receive voice input. Operation 443 can be understood by referring to the description of operation 423 of FIG. 4b. The electronic device (101) may receive a first voice input, for example, at a first time. For example, the electronic device (101) may receive the first voice input through a microphone (210) based on user speech.
[0140] In operation 444, according to one embodiment, the electronic device (101) can check whether a specified period has elapsed. For example, the electronic device (101) can check whether a subsequent voice input is received during a specified period (e.g., a first period) from the time when a voice input (e.g., a first voice input of operation 443) is received (e.g., a first time point of operation 443). For example, the electronic device (101) can check the length of time during which no voice input is received. According to one embodiment, the electronic device (101) can receive a subsequent voice input (e.g., a third voice input following the second voice input) based on checking that a subsequent voice input (e.g., a second voice input) is received within a specified period (e.g., a first period) from the time when a voice input (e.g., a first voice input of operation 443) is received (e.g., a first time point of operation 443).
[0141] In operation 445, according to one embodiment, the electronic device (101) may determine whether at least one condition for outputting a response (e.g., a response condition) is satisfied based on confirming that a subsequent voice input is not received for a specified period (e.g., a first period) from the time when a voice input (e.g., a first voice input of operation 443) is received (e.g., a first time point of operation 443). The at least one condition for outputting a response (e.g., a response condition) can be understood by referring to the description of operation 425 of FIG. 4b.
[0142] In operation 447, according to one embodiment, the electronic device (101) may output a response related to a voice input based on at least one of at least one condition for outputting a response (e.g., a response condition) being satisfied. For example, the electronic device (101) may output a response related to a voice input (e.g., a first voice input of operation 443) (e.g., a first response related to the first voice input) based on at least one of at least one of the conditions for outputting a response (e.g., a response condition) being satisfied, when a subsequent voice input is not received for a specified period (e.g., a first period) from the time when the voice input (e.g., a first voice input of operation 443) is received (e.g., a first time point of operation 443). The output of the response (e.g., a first response related to the first voice input) can be understood by referring to the description of operation 427 in FIG. 4b.
[0143] In operation 449, according to one embodiment, the electronic device (101) may receive a subsequent voice input (e.g., a second voice input following the first voice input) instead of outputting a response (e.g., a first response related to the first voice input) based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied. For example, the electronic device (101) may wait to receive a subsequent voice input (e.g., a second voice input following the first voice input) instead of outputting a response (e.g., a first response related to the first voice input) based on the fact that at least one condition for outputting a response (e.g., a response condition) is not satisfied, when a subsequent voice input is not received for a specified period (e.g., a first period) from the time when the voice input (e.g., the first voice input of operation 443) is received (e.g., a first time point of operation 443).
[0144] FIG. 4e is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 4e can be explained based on the embodiments of FIG. 1, 2a, 2b, 2c, and FIG. 3, the embodiments of FIG. 4a, 4b, 4c, and 4d, and the embodiments described below.
[0145] At least some of the operations of FIG. 4e may be omitted. The order of the operations of FIG. 4e may be changed. Operations other than those of FIG. 4e may be performed before, during, or after the operations of FIG. 4e.
[0146] FIG. 4e may be a diagram for explaining the operations of FIG. 4a in a chronological order. For example, the operations of FIG. 4e may be embodiments of the operations of FIG. 4a. Any parts of the description of the operations of FIG. 4e that overlap with the description of the operations of FIG. 4a may be omitted.
[0147] Referring to FIG. 4e, in operation 451, according to one embodiment, the electronic device (101) can continuously monitor the user's gaze toward the screen of the display (250). For example, the user can continuously look at the screen of the electronic device (101). For example, the electronic device (101) can continuously monitor the user's gaze toward the screen of the display (250). For example, the electronic device (101) can continuously monitor the user's gaze toward the screen of the display (250) from before receiving the first voice input until after receiving the first voice input.
[0148] In operation 453, according to one embodiment, the electronic device (101) can determine that no subsequent voice input is received for a specified period. For example, the electronic device (101) can determine that no subsequent voice input after the first voice input is received for a specified period while the user's continuous gaze toward the screen of the display (250) is being monitored. For example, the user may not speak for a specified period after inputting the first voice input while looking at the screen of the display (250).
[0149] In operation 455, according to one embodiment, the electronic device (101) may output a first response related to the first voice input based on confirming that a subsequent voice input is not received for a specified period from the time the first voice input is received, while confirming the continuous gaze of a user toward the screen of the display (250).
[0150] FIG. 4f is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 4f can be described based on the embodiments of FIG. 1, 2a, 2b, 2c and FIG. 3, the embodiments of FIG. 4a, 4b, 4c, 4d and 4e and the embodiments described below.
[0151] At least some of the operations of FIG. 4f may be omitted. The order of the operations of FIG. 4f may be changed. Operations other than those of FIG. 4f may be performed before, during, or after the operations of FIG. 4f.
[0152] Referring to FIG. 4f, in operation 461, according to one embodiment, the electronic device (101) may activate the camera (220) based on the activation of the conversation mode. According to one embodiment, the electronic device (101) may activate the camera (220) based on the activation of the conversation mode according to a setting (e.g., a first setting). According to one embodiment, the electronic device (101) may check whether an event causing the activation of the camera (220) occurs based on the activation of the conversation mode according to a setting (e.g., a second setting). According to one embodiment, even if the conversation mode is activated, the electronic device (101) may not activate the camera (220) based on the fact that other conditions are not met. Other conditions for activating the camera (220) will be described later.
[0153] 463 In operation, according to one embodiment, the electronic device (101) may disable the camera (220) based on the conversation mode being disabled. For example, the electronic device (101) may disable the camera (220) based on the conversation mode being terminated.
[0154] FIG. 5 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 5 can be explained based on the embodiments of FIG. 1, 2a, 2b, 2c and FIG. 3, the embodiments of FIG. 4a, 4b, 4c, 4d, 4e, 4f and the embodiments described below.
[0155] Referring to FIG. 5, according to one embodiment, the electronic device (101) can check the mounting state of the electronic device (101) and determine an operating mode based on different conditions depending on the mounting state.
[0156] At least some of the operations of FIG. 5 may be omitted. The order of the operations of FIG. 5 may be changed. Operations other than the operations of FIG. 5 may be performed before, during, or after the operations of FIG. 5.
[0157] Referring to FIG. 5, in operation 501, according to one embodiment, the electronic device (101) can determine the positioning state of the electronic device (101). For example, the electronic device (101) can determine the positioning state of the electronic device (101) by using sensing data (e.g., sensing value) of at least one sensor (e.g., the first sensor (231) of FIG. 2a) configured to identify the posture and / or movement of the electronic device. The positioning state of the electronic device (101) may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to a floor surface (e.g., the bottom surface of a desk), a standing state in which the back surface of the electronic device (101) is placed at a certain angle (e.g., a non-parallel angle) with respect to the floor surface (e.g., the bottom surface of a desk), and a handheld state in which the electronic device (101) is held by a user. For example, the electronic device (101) can determine the bottom state based on confirming that the back surface of the electronic device (101) is substantially parallel to the ground surface and that there is little to no movement of the electronic device (101). For example, in the bottom state of the electronic device (101), as the display (250) of the electronic device (101) faces upward (e.g., toward the sky), the user can visually perceive the display (250) of the electronic device (101). For example, the electronic device (101) can determine the standing state based on confirming that the back surface of the electronic device (101) forms a certain angle (e.g., not parallel) with the ground surface and that there is little to no movement of the electronic device (101). For example, the electronic device (101) can determine the gripping state based on detecting the movement of the electronic device (101). According to one embodiment, the electronic device (101) may determine an operation mode based on different conditions depending on the mounting state of the electronic device (101). The conditions depending on the mounting state may partially overlap.
[0158] In operation 503, according to one embodiment, the electronic device (101) may determine an operation mode based on a first condition regarding transition information in a floor state. The transition information may be information that causes the operation mode of the conversation mode to be switched. A condition based on other information verified based on voice data, in addition to the transition information, may be called the first condition. For example, at least one condition for outputting a response (e.g., a first response) (e.g., a response condition) may include an operation to check whether the voice input (e.g., voice data) satisfies the first condition regarding whether it includes transition information that causes the operation mode to be switched. The electronic device (101) may determine an operation mode based on other conditions (e.g., mouth movement, input via a button, biometric information) in addition to the first condition in a floor state, and this will be explained in the drawings described below. In operation 503, the electronic device (101) may check the first condition regarding whether the voice data includes transition information that causes the operation mode of the conversation mode to be switched in a floor state. According to one embodiment, the transition information may include information regarding a name. For example, the electronic device (101) may select an operation mode based on the identified name based on identifying the name in the voice data. An example of the name is described in detail in FIGS. 6 to 8. According to one embodiment, the switching information may include information about a filler voice. For example, the electronic device (101) may select an operation mode based on the identified filler voice based on identifying the filler voice in the voice data. An example of the filler voice is described in detail in FIG. 9.
[0159] In operation 505, according to one embodiment, the electronic device (101) may determine an operation mode based on a first condition (e.g., a first condition regarding transition information of operation 503) and a second condition regarding the direction of gaze while in a standing state. For example, the electronic device (101) may determine an operation mode based on the first condition and / or the second condition while in a standing state. For example, at least one condition for outputting a response (e.g., a first response) (e.g., a response condition) may include an operation to check whether the first condition and / or the second condition regarding the user's direction of gaze are satisfied. The electronic device (101) may determine an operation mode based not only on the first condition and / or the second condition but also on other conditions (e.g., mouth movement, input via a button, biometric information) while in a standing state, and this will be explained in the drawings described below. In operation 505, the electronic device (101) can check a first condition regarding whether voice data includes switching information that causes the voice data to switch the operation mode of the conversation mode while in a standing state. The electronic device (101) can determine the operation mode based on the first condition regarding the switching information while in a standing state. As described above in operation 503, the switching information may include information regarding the title or information regarding filler voice, which will be explained in detail in FIGS. 6 to 9. The electronic device (101) can check a second condition regarding the direction of gaze while in a standing state. For example, the electronic device (101) can check the user's direction of gaze using the camera (220) of the electronic device (101). The electronic device (101) can determine the operation mode based on the second condition regarding the direction of gaze while in a standing state. An example regarding the direction of gaze will be explained in detail in FIGS. 10 to 15. In addition to the direction of gaze, other information confirmed based on the image data of the camera (220) can be called a second condition.For example, the electronic device (101) can check the user's gestures and / or facial expressions using the camera (220) of the electronic device (101). For example, the electronic device (101) can determine an operation mode based on a condition (e.g., a second condition) regarding the gestures and / or facial expressions while in a standing state. Examples of gestures and / or facial expressions are described in detail in FIGS. 16 to 20.
[0160] In operation 507, according to one embodiment, the electronic device (101) may determine an operation mode based on a first condition (e.g., a first condition regarding transition information of operation 503), a second condition (e.g., a second condition regarding gaze direction, gesture, and / or facial expression of operation 505), and a third condition regarding distance while in a gripping state. For example, the electronic device (101) may determine an operation mode based on the first condition, the second condition, and / or the third condition while in a gripping state. For example, at least one condition for outputting a response (e.g., a first response) (e.g., a response condition) may include an operation to check whether the first condition, the second condition, and / or the third condition regarding the first distance between the electronic device (101) and the user is satisfied. The electronic device (101) may determine an operation mode based on a first condition, a second condition, and / or a third condition, as well as other conditions (e.g., mouth movement, input via a button, biometric information), which will be explained in the drawings described below. In operation 507, the electronic device (101) may check a first condition regarding whether voice data includes transition information that causes the operation mode of the conversation mode to be switched while in a gripping state. The electronic device (101) may determine an operation mode based on the first condition regarding transition information while in a gripping state. As described above in operation 503, the transition information may include information regarding the appellation or information regarding filler voice, which will be explained in detail in FIGS. 6 through 9. The electronic device (101) may check a second condition (e.g., a condition regarding gaze direction, gesture, and / or facial expression) while in a gripping state. The electronic device (101) can determine a mode of operation based on a second condition (e.g., a condition regarding gaze direction, gesture, and / or facial expression) while in a gripping state.As described above in operation 505, the second condition may include conditions regarding gaze direction, gesture, and / or facial expression, which will be explained in detail in FIGS. 10 to 20. The electronic device (101) can determine a third condition regarding distance while in a gripping state. For example, the electronic device (101) can determine the distance between the electronic device (101) and the user. For example, the electronic device (101) can determine the distance between the electronic device (101) and the user by using sensing data (e.g., sensing value) of at least one sensor (e.g., the second sensor (232) in FIG. 2a). For example, the electronic device (101) can determine the distance between the electronic device (101) and the user by using a microphone (210) (e.g., a plurality of microphones). The electronic device (101) can determine an operation mode based on the third condition regarding the distance between the electronic device (101) and the user while in a gripping state. An example of the distance will be described in detail in FIGS. 21 and FIGS. 22. A condition based on the distance between an external device (e.g., the external device (3310, 3320, 3330) of FIG. 33) (e.g., a wearable device)) and the user, as well as the distance between the electronic device (101) and the user, can be referred to as a third condition. For example, the electronic device (101) can determine the distance between the external device (e.g., 3310, 3320, 3330) of FIG. 33) (e.g., a wearable device) and the user. For example, the electronic device (101) can determine an operation mode based on the condition based on the distance between the external device (e.g., 3310, 3320, 3330) of FIG. 33) (e.g., a wearable device) and the user (e.g., the third condition) while in a gripping state. An example regarding the distance between the external device (e.g., 3310, 3320, 3330) of FIG. 33) (e.g., a wearable device) and the user is FIG. We will explain this in detail in section 33.
[0161] FIG. 6 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 6 can be described based on the embodiments of FIG. 1 to 5, the embodiments of FIG. 7 and 8, and the embodiments described below.
[0162] Referring to FIG. 6, according to one embodiment, the electronic device (101) can select an operation mode based on the identified name based on confirming the name in voice data.
[0163] At least some of the operations of FIG. 6 may be omitted. The order of the operations of FIG. 6 may be changed. Operations other than those of FIG. 6 may be performed before, during, or after the operations of FIG. 6.
[0164] Referring to FIG. 6, in operation 601, according to one embodiment, the electronic device (101) can determine whether voice data includes switching information that causes the operation mode of the conversation mode to be switched. The switching information may be information that causes the operation mode of the conversation mode to be switched. For example, the switching information may include information about a title. The electronic device (101) can determine whether the voice data includes a title.
[0165] In operation 603, according to one embodiment, the electronic device (101) can identify a designation in voice data. The designation may include a designation assigned to an artificial intelligence model (e.g., voice agent) and / or a designation assigned to a user. The electronic device (101) can determine whether a designation included in the voice data is a designation assigned to an artificial intelligence model (e.g., voice agent). The electronic device (101) can determine whether a designation included in the voice data is a designation assigned to a user. A designation assigned to an artificial intelligence model (e.g., voice agent) may include a default designation assigned to the artificial intelligence model (e.g., voice agent) and / or a designation assigned to the artificial intelligence model (e.g., voice agent) by a user. A designation assigned to a user may include a default designation assigned to a user and / or a designation assigned to a user by a user. For example, the user may input information regarding a title to be assigned to an artificial intelligence model (e.g., a voice agent) and / or information regarding a title to be assigned to the user. The electronic device (101) may verify, based on the input, information regarding a title to be assigned to an artificial intelligence model (e.g., a voice agent) and / or information regarding a title to be assigned to the user. If information regarding a title to be assigned to an artificial intelligence model (e.g., a voice agent) is not input, the electronic device (101) may assign a default title corresponding to the artificial intelligence model (e.g., a voice agent) to the artificial intelligence model (e.g., a voice agent). If information regarding a title to be assigned to the user is not input, the electronic device (101) may assign a default title corresponding to the user to the user. There are no restrictions on the default title (e.g., a default title corresponding to the artificial intelligence model (e.g., a voice agent) or a default title corresponding to the user). The user may input information regarding the titles of others as well as their own titles, which will be described later.According to one embodiment, the electronic device (101) may perform an operation to receive information about a title before entering a conversation mode. According to one embodiment, the electronic device (101) may perform an operation to receive information about a title after entering a conversation mode. According to one embodiment, information about a title may be changed while the conversation mode is being performed.
[0166] In operation 605, according to one embodiment, the electronic device (101) may operate in a response mode based on identifying a designation that causes a response mode in voice data. For example, at least one condition (e.g., response condition) for outputting a response (e.g., first response) may include an operation of identifying a designation assigned to an artificial intelligence model (e.g., voice agent) in a voice input (e.g., first voice input). For example, the electronic device (101) may select the response mode as an operation mode based on identifying a designation assigned to an artificial intelligence model (e.g., voice agent) in voice data. According to one embodiment, the electronic device (101) may select the response mode as an operation mode based on identifying a user's utterance requesting an operation of an artificial intelligence model (e.g., voice agent) in voice data according to an analysis of voice data, even if it does not identify a designation assigned to an artificial intelligence model (e.g., voice agent) in voice data.
[0167] In operation 607, according to one embodiment, the electronic device (101) may operate in a listening mode based on identifying a designation that causes a listening mode in voice data. For example, the electronic device (101) may select the listening mode as an operating mode based on identifying a designation assigned to a user in voice data. According to one embodiment, even if the electronic device (101) does not identify a designation assigned to a user in voice data, it may select the listening mode as an operating mode based on identifying the absence of a user's utterance requesting the operation of an artificial intelligence model (e.g., voice agent) in voice data according to an analysis of voice data.
[0168] FIG. 7 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 7 may be described based on the embodiments of FIG. 1 to 6, the embodiments of FIG. 8, and embodiments described below. FIG. 8 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0169] Referring to FIG. 7, according to one embodiment, the electronic device (101) can select an operation mode based on a designation in group mode.
[0170] At least some of the operations of FIG. 7 may be omitted. The order of the operations of FIG. 7 may be changed. Operations other than the operations of FIG. 7 may be performed before, during, or after the operations of FIG. 7.
[0171] Referring to FIG. 7, in operation 701, according to one embodiment, the electronic device (101) may enter a group mode. The group mode may be one of the operation modes of the conversation mode. The group mode may be a mode in which multiple users participate in a conversation. The opposite concept of the group mode may be the basic mode. The basic mode may be a mode in which one user participates in a conversation. The electronic device (101) may enter the group mode based on confirming an event that causes entry into the group mode in which multiple users participate in a conversation. For example, referring to FIG. 8(a), the electronic device (101) may enter the group mode based on a call to the group mode. The call to the group mode may be a voice command that causes the electronic device (101) to enter the group mode. For example, the electronic device (101) may enter the group mode based on receiving voice data corresponding to an utterance that includes a voice that causes the group mode to be triggered (e.g., “Bixby! I’m going to group talk~”). According to one embodiment, the electronic device (101) may enter group mode when a primary user wishes to converse with multiple users. According to one embodiment, the primary user may be a user registered as the primary user of the electronic device (101). According to one embodiment, the primary user may be the user with the largest face area in an image obtained through the camera (220). For example, the electronic device (101) may enter group mode based on identifying the primary user and confirming a voice that triggers group mode (e.g., “Bixby! I’ll have a group conversation~”) in voice data corresponding to the primary user’s speech. For example, the electronic device (101) may identify the user registered as the primary user based on voice analysis of voice data. For example, the electronic device (101) may identify the user registered as the primary user based on face analysis of image data.For example, the electronic device (101) can determine the primary user by comparing the size of the face region in the image data. According to one embodiment, the electronic device (101) may enter group mode based on receiving voice data corresponding to an utterance that triggers group mode (e.g., “Bixby! I’m going to group chat~”) without distinguishing between the primary user and other users. According to one embodiment, the electronic device (101) may enter group mode based on confirming the voices of multiple users in the voice data. For example, the electronic device (101) may enter group mode based on confirming the voices of multiple users in the voice data without calling for group mode.
[0172] In operation 703, according to one embodiment, the electronic device (101) may perform a procedure for registering participants in group mode. For example, the electronic device (101) may check input for titles corresponding to multiple users so as to assign titles corresponding to multiple users respectively in group mode. For example, in FIG. 8(b), the electronic device (101) may display a screen for registering participants to participate in a group conversation. For example, while the screen for registering participants to participate in a group conversation is displayed so as to assign titles corresponding to multiple users respectively, the main user may input information for titles corresponding to multiple users into the electronic device (101). Based on receiving information for titles corresponding to multiple users, the electronic device (101) may assign titles corresponding to each of the multiple users. According to one embodiment, the electronic device (101) can assign titles corresponding to each of a plurality of users based on confirming voice for registering titles corresponding to a plurality of users in voice data, without displaying a screen for registering participants to participate in a group conversation. For example, a user may say, "Register Jack and Michel," and the electronic device (101) can assign titles corresponding to each of a plurality of users based on voice data. In FIG. 8(b), the electronic device (101) may receive information regarding a title corresponding to an artificial intelligence model (e.g., voice agent) in group mode. If information regarding a title to be assigned to an artificial intelligence model (e.g., voice agent) is not received, the electronic device (101) may assign a default title corresponding to the artificial intelligence model (e.g., voice agent) to the artificial intelligence model (e.g., voice agent). The electronic device (101) may receive information regarding a title corresponding to an artificial intelligence model (e.g., voice agent) in basic mode.
[0173] In operation 705, according to one embodiment, the electronic device (101) can identify a designation in voice data while operating in group mode. The designation may include a designation assigned to an artificial intelligence model (e.g., voice agent) and / or a designation assigned to a plurality of users.
[0174] In operation 707, according to one embodiment, the electronic device (101) may operate in a response mode based on identifying a designation that causes a response mode in voice data in a group mode. For example, at least one condition (e.g., response condition) for outputting a response (e.g., first response) may include an operation of identifying a designation assigned to an artificial intelligence model (e.g., voice agent) in a voice input (e.g., first voice input). For example, the electronic device (101) may select the response mode as an operation mode based on identifying a designation assigned to an artificial intelligence model (e.g., voice agent) in voice data in a group mode. According to one embodiment, the electronic device (101) may select the response mode as an operation mode based on identifying a user's utterance requesting an operation of an artificial intelligence model (e.g., voice agent) in voice data according to an analysis of voice data, even if it does not identify a designation assigned to an artificial intelligence model (e.g., voice agent) in voice data.
[0175] In operation 709, according to one embodiment, the electronic device (101) may operate in a listening mode based on identifying a designation that causes a listening mode in voice data in group mode. For example, the electronic device (101) may select the listening mode as an operating mode based on identifying at least one of the designations assigned to multiple users in voice data in group mode. For example, the electronic device (101) may receive a subsequent voice input instead of outputting a response based on identifying at least one of the designations corresponding to multiple users in voice input in group mode.
[0176] According to one embodiment, the electronic device (101) may select a listening mode as an operating mode based on confirming the absence of a user’s utterance requesting the operation of an artificial intelligence model (e.g., voice agent) in the voice data according to an analysis of the voice data, even if it cannot confirm a designation assigned to the user in the voice data.
[0177] FIG. 9 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 9 can be described based on the embodiments of FIG. 1 to 8 and the embodiments described below.
[0178] Referring to FIG. 9, according to one embodiment, the electronic device (101) can select an operation mode based on the identified filler voice (e.g., "Hmm...", "Let's wait a moment") in voice data.
[0179] At least some of the operations of FIG. 9 may be omitted. The order of the operations of FIG. 9 may be changed. Operations other than those of FIG. 9 may be performed before, during, or after the operations of FIG. 9.
[0180] Referring to FIG. 9, in operation 901, according to one embodiment, the electronic device (101) can determine whether voice data includes switching information that causes the operation mode of the conversation mode to be switched. The switching information may be information that causes the operation mode of the conversation mode to be switched. For example, the switching information may include information about filler voice.
[0181] In operation 903, according to one embodiment, the electronic device (101) can determine whether the voice data includes filler voice. The filler voice (e.g., "Hmm...", "Let's wait a moment.") may be a voice for a word or sentence that serves to buy time to think or to connect speech. There is no limitation on the type of filler voice.
[0182] In operation 905, according to one embodiment, the electronic device (101) may select the listening mode as the operating mode based on identifying a filler voice (e.g., "Hmm...", "Let's wait a moment") that causes the listening mode in voice data.
[0183] In operation 907, according to one embodiment, the electronic device (101) may select a response mode as an operation mode based on other conditions, based on confirming the absence of filler voice that causes a listening mode in voice data.
[0184] FIG. 10 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 10 may be described based on the embodiments of FIG. 1 to 9, the embodiments of FIG. 11 to 15, and the embodiments described below.
[0185] FIG. 11 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0186] Referring to FIG. 10, according to one embodiment, the electronic device (101) can select an operation mode according to the user's line of sight.
[0187] At least some of the operations of FIG. 10 may be omitted. The order of the operations of FIG. 10 may be changed. Operations other than those of FIG. 10 may be performed before, during, or after the operations of FIG. 10.
[0188] Referring to FIG. 10, in operation 1001, according to one embodiment, the electronic device (101) can determine the user's gaze. For example, the electronic device (101) can determine the direction of the user's gaze by using the camera (220) of the electronic device (101). For example, in FIG. 11, the electronic device (101) can set a specific point (e.g., endpoint) of the user's (1100) eye in an image obtained using the camera (220) as a tracking point (e.g., 1110). The electronic device (101) can determine the user's (1100) gaze by tracking the location of the tracking point (e.g., 1110) corresponding to the user's (1100) eye.
[0189] In operation 1003, according to one embodiment, the electronic device (101) can check the direction of the user's gaze and determine the operation mode of the conversation mode based on the direction of the gaze. For example, the electronic device (101) may check only the gaze of the main user or check the gazes of multiple users participating in the conversation. The operation of checking the main user can be understood by referring to the description of operation 701 of FIG. 7.
[0190] In operation 1005, according to one embodiment, the electronic device (101) may select a response mode as an operation mode based on confirming a first direction of gaze toward the electronic device (101) (e.g., the screen of the display (250)). For example, at least one condition (e.g., response condition) for outputting a response (e.g., the first response) may include an operation of confirming a first direction of gaze toward the electronic device (101) (e.g., the screen of the display (250)). For example, the electronic device (101) may select a response mode as an operation mode based on the fact that the direction of the gaze of the main user is a first direction toward the electronic device (101) (e.g., the screen of the display (250)). For example, the electronic device (101) may select a response mode as an operation mode based on the fact that the direction of gaze of not only the main user but also other users participating in the conversation is a first direction toward the electronic device (101) (e.g., the screen of the display (250)). According to one embodiment, the electronic device (101) may ignore the gaze of another user based on the fact that the direction of the gaze of another user, other than the gaze of the main user, is a first direction toward the electronic device (101) (e.g., the screen of the display (250)). According to one embodiment, the electronic device (101) may ignore the other gaze based on confirming the other gaze of a person other than the user participating in the conversation.
[0191] In operation 1007, according to one embodiment, the electronic device (101) may select a listening mode as an operating mode based on not confirming a gaze in a first direction (e.g., a direction toward the electronic device (101) (e.g., a screen of the display (250)). The electronic device (101) may select a listening mode as an operating mode based on confirming a gaze in a direction different from the first direction (e.g., a direction toward the electronic device (101) (e.g., a screen of the display (250)). According to one embodiment, the electronic device (101) may select a response mode as an operating mode based on other conditions even if a gaze in a direction different from the first direction (e.g., a direction toward the electronic device (101) (e.g., a screen of the display (250)) is confirmed.
[0192] FIG. 12 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure.
[0193] Referring to FIG. 12, in operation 1201, according to one embodiment, an electronic device (101) may activate a camera (220). The electronic device (101) may activate the camera (220) to acquire image data. For example, the electronic device (101) may activate the camera (220) of the electronic device (101) based on checking a standing state or a gripping state. The electronic device (101) may activate the camera (220) to acquire image data in a standing state or a gripping state. For example, the electronic device (101) may not activate the camera (220) in a floor state.
[0194] In operation 1203, according to one embodiment, the electronic device (101) may perform calibration of the camera (220). For example, the electronic device (101) may perform calibration of the camera (220) in a standing state or in a holding state. Calibration may be an operation to detect the position of the user's eyes and test to track a tracking point corresponding to the eyes in image data in order to track gaze through the camera (220). For example, the electronic device (101) may perform calibration by detecting the user's eyes, setting a tracking point corresponding to the eyes, displaying an object (e.g., a point displayed on the screen) at a reference point on the screen, and checking the movement of the tracking point when the user looks at the object. The electronic device (101) may also correct the calibration result by displaying an object (e.g., a point displayed on the screen) at an additional reference point to test accuracy after calibration is performed, and checking the movement of the tracking point when the user looks at the object. For example, in a mounted state, calibration can be easily performed because there is little or no movement of the camera (220), but in a held state, calibration of the camera (220) may not be completed due to the movement of the electronic device (101). In the case where calibration is impossible (e.g., when tracking of the gaze direction fails), the electronic device (101) may indicate that the eye is not recognized, disable the camera (220), and determine an operation mode based on the remaining conditions excluding the conditions regarding the gaze.
[0195] FIG. 13 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0196] For example, in FIG. 13(a), the electronic device (101) may enter conversation mode based on acknowledging a user's call (e.g., "Bixby!") while in a standing state. As the electronic device (101) enters conversation mode while in a standing state, it may perform calibration of the camera (220). For example, in FIG. 13(b), the electronic device (101) may display a screen (1310) containing a guidance message (e.g., 1311) to look at the camera (220). For example, the electronic device (101) may display an object (e.g., 1312) that guides the position of the camera (220). For example, the electronic device (101) may apply an effect that guides the position of the camera (220) (e.g., an effect applied to the area surrounding the camera (220) on the screen). Accordingly, in FIG. 13(b), the user may look at the camera (220). Subsequently, the electronic device (101) may perform calibration and display a screen (1320) containing a phrase (e.g., 1321) indicating the completion of calibration in FIG. 13 (c). For example, the electronic device (101) may apply an effect indicating the completion of calibration. For example, the effect indicating the completion of calibration may be an effect applied to the area surrounding the camera (220) in the screen (1320), and may be different from the effect in FIG. 13 (b). For example, the color of the effect in FIG. 13 (b) and the color of the effect in FIG. 13 (c) may be different. The description of the effect is exemplary, and those skilled in the art will understand that various effects may be applied.
[0197] 1205 In operation, according to one embodiment, the electronic device (101) may track the user's gaze based on the completion of calibration and determine an operation mode based on the direction of the gaze. The determination of an operation mode based on the direction of the gaze can be understood by referring to the embodiments described above.
[0198] According to one embodiment, the electronic device (101) may disable the camera (220) based on a transition to a floor state, failure of tracking the line of sight, or termination of the conversation mode.
[0199] FIG. 14 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0200] FIG. 15 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0201] Referring to Figures 14 and 15, the conditions for the line of sight will be explained.
[0202] For example, referring to FIG. 14, the user may look at or not look at the electronic device (101) while performing English conversation practice. For example, in FIG. 14 (a), the user may practice English conversation using the electronic device (101). According to one embodiment, the electronic device (101) may select a listening mode as an operating mode based on confirming the user's gaze not looking at the electronic device (101) while the user is performing English conversation practice, for example, as in FIG. 14 (b). According to one embodiment, the electronic device (101) may select a response mode as an operating mode based on confirming the user's gaze looking at the electronic device (101) while the user is performing English conversation practice, for example, as in FIG. 14 (c). Accordingly, the user may control the operating mode of the conversation mode of the electronic device (101) through the action of looking at or not looking at the electronic device (101).
[0203] For example, referring to FIG. 15, a first user (1510) and a second user (1520) can converse around an electronic device (101). In FIG. 15 (a), the electronic device (101) can select a listening mode as an operating mode based on not being able to confirm the gaze of the first user (1510) and the second user (1520) looking at the electronic device (101). The electronic device (101) can select a listening mode as an operating mode based on confirming the gaze of the first user (1510) looking in a direction other than the electronic device (101) and the gaze of the second user (1520) looking in a direction other than the electronic device (101). In FIG. 15 (a), the electronic device (101) can store voice data acquired during the listening mode in a memory (130). Subsequently, in FIG. 15 (b), the first user (1510) may look at the electronic device (101). The electronic device (101) may select a response mode as an operation mode based on confirming the gaze of the first user (1510) looking at the electronic device (101). In FIG. 15 (c), the electronic device (101) may check response information generated based on voice data (e.g., "Find me a good highball place nearby") and output a response (e.g., "XXX Sangsu Point OOO 38°C is famous.") based on the confirmed response information. According to one embodiment, even if no separate voice data is acquired after switching to the response mode, the electronic device (101) may check response information generated based on voice data accumulated while operating in conversation mode (e.g., voice data stored during the listening mode of FIG. 15 (a)) and output a response (e.g., "XXX Sangsu Point OOO 38°C is famous.") based on the confirmed response information.
[0204] FIG. 16 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 16 may be described based on the embodiments of FIG. 1 to 15, the embodiments of FIG. 17 to 20, and the embodiments described below. FIG. 17 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 18 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 19 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 20 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0205] Referring to FIG. 16, according to one embodiment, the electronic device (101) can select an operation mode according to a user's gesture.
[0206] At least some of the operations of FIG. 16 may be omitted. The order of the operations of FIG. 16 may be changed. Operations other than those of FIG. 16 may be performed before, during, or after the operations of FIG. 16.
[0207] Referring to FIG. 16, in operation 1601, according to one embodiment, an electronic device (101) can identify a user's gesture. For example, the electronic device (101) can identify a user's gesture by using a camera (220) of the electronic device (101). For example, the electronic device (101) can identify feature points corresponding to the user's body in an image obtained using the camera (220), and identify the user's gesture based on a combination of feature points corresponding to the body and the movement of the feature points. According to one embodiment, the electronic device (101) can also identify a user's facial expression. For example, the electronic device (101) can identify feature points corresponding to the user's face in an image obtained using the camera (220), and identify the user's facial expression based on a combination of feature points corresponding to the face and the movement of the feature points.
[0208] In operation 1603, according to one embodiment, the electronic device (101) can check for the presence or absence of a gesture (or facial expression) that causes deactivation.
[0209] In operation 1605, according to one embodiment, the electronic device (101) may terminate or temporarily suspend the conversation mode based on identifying a gesture (or facial expression) that causes deactivation in an image obtained using the camera (220). The electronic device (101) may disable the camera (220) based on identifying a gesture (or facial expression) that causes deactivation. For example, in FIG. 17 (a), the electronic device (101) may register a first gesture (1710) (e.g., a gesture of opening the hand to show the palm) as a gesture that causes deactivation. The electronic device (101) may terminate or temporarily suspend the conversation mode based on identifying a gesture that causes deactivation (e.g., a gesture of opening the hand to show the palm). According to one embodiment, the electronic device (101) may terminate or temporarily suspend the conversation mode based on whether the gesture being observed is a first gesture (e.g., a gesture of opening the hand to show the palm) while the first direction of sight toward the electronic device (101) (e.g., the screen of the display (250)) is not observed.
[0210] 1607 In operation, according to one embodiment, the electronic device (101) may check for the presence or absence of a gesture (or facial expression) that causes a switch in the operation mode. According to one embodiment, the electronic device (101) may switch the operation mode based on checking for a gesture that causes a switch in the operation mode. According to one embodiment, the gesture that causes a switch from listening mode to response mode and the gesture that causes a switch from response mode to listening mode may be different.
[0211] In operation 1609, according to one embodiment, the electronic device (101) may select the listening mode as an operating mode based on the identification of a gesture that causes the listening mode in an image acquired using the camera (220). According to one embodiment, the electronic device (101) may select the listening mode as an operating mode based on the identification of a gesture that causes the listening mode while a first direction of gaze toward the electronic device (101) (e.g., the screen of the display (250)) is not identified. For example, the electronic device (101) may receive a subsequent voice input instead of outputting a response based on the identification of a gesture that causes the listening mode. For example, the electronic device (101) may receive a subsequent voice input instead of outputting a response based on the identification of a gesture that causes the listening mode while the user's gaze is not identified. For example, in FIG. 17(b), the electronic device (101) can select the listening mode as the operating mode based on confirming a second gesture (1720) (e.g., a thumbs up) that causes the listening mode.
[0212] In operation 1611, according to one embodiment, the electronic device (101) may select a response mode as an operation mode based on the identification of a gesture that triggers a response mode in an image acquired using a camera (220). For example, at least one condition (e.g., response condition) for outputting a response (e.g., a first response) may include an operation of identifying a gesture that triggers a response. According to one embodiment, the electronic device (101) may select a response mode as an operation mode based on the identification of a gesture that triggers a response mode while a first direction of gaze toward the electronic device (101) (e.g., a screen of a display (250)) is not identified. For example, the electronic device (101) may output a response related to voice input based on the identification of a gesture that triggers a response. For example, the electronic device (101) may output a response related to voice input based on the identification of a gesture that triggers a response while the user's gaze is not identified. For example, in (c) of FIG. 17, the electronic device (101) can select the response mode as the operation mode based on confirming a third gesture (1730) (e.g., nodding) that causes the response mode.
[0213] According to one embodiment, the gesture causing the transition from listening mode to response mode and the gesture causing the transition from response mode to listening mode may be the same. For example, the electronic device (101) may transition from listening mode to response mode based on confirming the gesture causing the transition of the operation mode. Subsequently, the electronic device (101) may transition from response mode to listening mode based on confirming the same gesture causing the transition of the operation mode. As the gesture causing the transition of the operation mode is repeated, the electronic device (101) may continue to transition from listening mode to response mode, or from response mode to listening mode.
[0214] According to one embodiment, the electronic device (101) can learn the user's gestures, facial expressions, and gaze. For example, the electronic device (101) can perform learning using an artificial intelligence model (e.g., a voice agent) based on data regarding the user's gestures, facial expressions, and gaze, and data regarding the user's intention. The electronic device (101) can identify the user's intention corresponding to at least one combination of the user's gestures, facial expressions, or gaze. The electronic device (101) can infer the user's intention using an artificial intelligence model (e.g., a voice agent) by using information regarding the user's gestures, facial expressions, and gaze as input data. The electronic device (101) can identify information regarding the user's gestures, facial expressions, and gaze in image data, infer the user's intention based on the identified information regarding the gestures, facial expressions, and gaze, and determine a listening mode or a response mode as an action mode based on the user's intention.
[0215] Referring to FIGS. 18 and 19, gestures that cause a pause or listening mode in conversation mode and gestures that cause a reactivation or response mode in conversation mode can be described.
[0216] In FIG. 18(a), according to one embodiment, the electronic device (101) may pause the conversation mode or operate in a listening mode based on detecting a first gesture (1810) (e.g., clenching a fist) that causes the conversation mode to pause or the listening mode. For example, the electronic device (101) may detect the first gesture (1810) (e.g., clenching a fist) based on image data and pause the conversation mode or operate in a listening mode.
[0217] In FIG. 18(b), according to one embodiment, the electronic device (101) may reactivate the conversation mode or operate in the response mode based on detecting a second gesture (1820) (e.g., opening a fist) that causes the reactivation of the conversation mode or the response mode. For example, the electronic device (101) may detect the second gesture (1820) (e.g., opening a fist) based on image data and reactivate the conversation mode or operate in the response mode.
[0218] In FIG. 19(a), according to one embodiment, the electronic device (101) may pause the conversation mode or operate in a listening mode based on confirming a first gesture (1910) (e.g., a fist clenched) that causes the conversation mode to pause or the listening mode to listen. For example, the electronic device (101) may receive a signal indicating that the first gesture (1910) (e.g., a fist clenched) is confirmed from a wearable device (1900) worn on the user's wrist and may pause the conversation mode or operate in a listening mode. For example, the wearable device (1900) may confirm the user's first gesture (1910) (e.g., a fist clenched) based on a biosignal and transmit a signal indicating that the first gesture (1910) (e.g., a fist clenched) is confirmed to the electronic device (101).
[0219] In FIG. 19(b), according to one embodiment, the electronic device (101) may reactivate the conversation mode or operate in the response mode based on confirming a second gesture (1920) (e.g., fist open) that causes the reactivation of the conversation mode or the response mode. For example, the electronic device (101) may receive a signal indicating that the second gesture (1920) (e.g., fist open) is confirmed from a wearable device (1900) worn on the user's wrist and may reactivate the conversation mode or operate in the response mode. For example, the wearable device (1900) may confirm the user's second gesture (1920) (e.g., fist open) based on a biosignal and transmit a signal indicating that the second gesture (1920) (e.g., fist open) is confirmed to the electronic device (101).
[0220] Referring to Fig. 20, the conditions for the gesture will be explained.
[0221] For example, referring to FIG. 20, a user can control the electronic device (101) through gestures without looking at the electronic device (101) while performing English conversation practice. For example, in FIG. 20 (a), a user can practice English conversation using the electronic device (101). According to one embodiment, in FIG. 20 (b), the electronic device (101) can select a listening mode as an operating mode based on confirming the user's gaze not looking at the electronic device (101) while the user is performing English conversation practice. According to one embodiment, in FIG. 20 (c), the electronic device (101) can select a response mode as an operating mode based on confirming a gesture (e.g., nodding) that triggers a response mode, even though the user is performing English conversation practice and the user is not looking at the electronic device (101).
[0222] FIG. 21 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 21 may be described based on the embodiments of FIG. 1 to 20, the embodiments of FIG. 22, and embodiments described below. FIG. 22 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0223] Referring to FIG. 21, according to one embodiment, the electronic device (101) may select an operation mode based on the distance between the electronic device (101) and the user. According to one embodiment, the electronic device (101) may also select an operation mode based on the distance between an external device (e.g., the external device (3310, 3320, 3330) of FIG. 33) and the user.
[0224] At least some of the operations of FIG. 21 may be omitted. The order of the operations of FIG. 21 may be changed. Operations other than those of FIG. 21 may be performed before, during, or after the operations of FIG. 21.
[0225] Referring to FIG. 21, in operation 2101, according to one embodiment, the electronic device (101) can determine the distance between the electronic device (101) and the user. For example, the electronic device (101) can determine the distance between the electronic device (101) and the user based on at least one sensor (e.g., the second sensor (232) of FIG. 2a). For example, the electronic device (101) can determine the distance between the electronic device (101) and the user by using a microphone (210) (e.g., a plurality of microphones). The operation of determining the distance has been described above. According to one embodiment, the electronic device (101) can perform an operation of measuring the distance between the electronic device (101) and the user based on determining the gripping state.
[0226] In operation 2103, according to one embodiment, the electronic device (101) can determine whether the electronic device (101) and the user are close by comparing the distance between the electronic device (101) and the user with a reference distance.
[0227] In operation 2105, according to one embodiment, the electronic device (101) may select a listening mode as an operating mode based on the distance between the electronic device (101) and the user being less than a reference distance. For example, in FIG. 22, the user (2200) may control the electronic device (101) to operate in a listening mode by holding the electronic device (101) with a hand (2210) and bringing the electronic device (101) near the mouth (2220).
[0228] In operation 2107, according to one embodiment, the electronic device (101) may select a response mode as an operation mode based on the distance between the electronic device (101) and the user being greater than or equal to a reference distance. For example, at least one condition (e.g., response condition) for outputting a response (e.g., a first response) may include an operation of confirming that the distance between the electronic device (101) and the user is greater than or equal to a reference distance. For example, the user may control the electronic device (101) to operate in a response mode by positioning the electronic device (101) away from the user while holding the electronic device (101) in their hand.
[0229] According to one embodiment, the electronic device (101) may stop measuring the distance between the electronic device (101) and the user based on a transition to a floor state or a standing state, or the end of a conversation mode.
[0230] FIG. 23 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 23 may be described based on the embodiments of FIG. 1 to 22, the embodiments of FIG. 24, and embodiments described below. FIG. 24 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0231] Referring to FIG. 23, according to one embodiment, the electronic device (101) can select an operation mode according to the shape of the user's mouth.
[0232] At least some of the operations of FIG. 23 may be omitted. The order of the operations of FIG. 23 may be changed. Operations other than those of FIG. 23 may be performed before, during, or after the operations of FIG. 23.
[0233] Referring to FIG. 23, in operation 2301, according to one embodiment, the electronic device (101) can detect the movement of the user's mouth. For example, in FIG. 24, the electronic device (101) can set a point corresponding to the user's (2400) mouth in an image obtained using the camera (220) of the electronic device (101) as a tracking point (2410). The electronic device (101) can detect the movement of the user's (2400) mouth by tracking the location of the tracking point (2410) corresponding to the user's (2400) mouth.
[0234] In operation 2303, according to one embodiment, the electronic device (101) may ignore voice data that is detected while there is no movement of the user's mouth. The electronic device (101) may ignore voice data that is detected while the user's mouth movement is not detected. For example, the electronic device (101) may not store voice data that is detected while the user's mouth movement is not detected in memory (130). For example, the electronic device (101) may ignore transition information even if transition information is detected in voice data that is detected while the user's mouth movement is not detected.
[0235] FIG. 25 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 25 may be described based on the embodiments of FIG. 1 to 24, the embodiments of FIG. 26, and embodiments described below. FIG. 26 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0236] Referring to FIG. 25, according to one embodiment, the electronic device (101) can select an operation mode based on input through a button.
[0237] At least some of the operations of FIG. 25 may be omitted. The order of the operations of FIG. 25 may be changed. Operations other than those of FIG. 25 may be performed before, during, or after the operations of FIG. 25.
[0238] Referring to FIG. 25, in operation 2501, according to one embodiment, the electronic device (101) can receive input through a button. The button may be a hardware button of the electronic device (101) or a software button displayed on the screen of the display (250) of the electronic device (101). For example, a user may press a hardware button or touch a software button.
[0239] In operation 2503, according to one embodiment, the electronic device (101) may determine an operation mode based on input via a button. According to one embodiment, the electronic device (101) may determine an operation mode based on other conditions based on the fact that input via a button is not acknowledged. According to one embodiment, the electronic device (101) may select a response mode as an operation mode based on the fact that input via a button is acknowledged. For example, at least one condition for outputting a response (e.g., a response condition) may include an operation of acknowledging input via a button. For example, the electronic device (101) may output a response related to voice input based on acknowledging input via a button. For example, the electronic device (101) may operate in a response mode as the user speaks while pressing (or touching) the button, or as the user speaks after pressing (or touching) the button. For example, in FIG. 26, a first user (2610) and a second user (2620) may be located around the electronic device (101). In FIG. 26 (a), the electronic device (101) may determine an operation mode based on other conditions, based on the fact that input via a button is not confirmed. For example, in FIG. 26 (a), the electronic device (101) may operate in a listening mode based on other conditions. In FIG. 26 (a), the electronic device (101) may store voice data acquired while operating in the listening mode in memory (130). Subsequently, in FIG. 26 (b), the electronic device (101) may operate in a response mode based on the fact that input via a button is confirmed. For example, in FIG. 26 (b), the user may speak "What do you think?" while pressing the button.The electronic device (101) operates in response mode based on confirmation of input through a button, and can output a response (e.g., sound) using voice data accumulated during conversation mode based on confirmation of an utterance requesting the operation of the electronic device (101) in voice data.
[0240] FIG. 27 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 27 may be described based on the embodiments of FIG. 1 to 26 and the embodiments described below.
[0241] In FIG. 27 (a), according to one embodiment, the electronic device (101) may enter conversation mode based on confirming a call to the electronic device (101) (e.g., “Hi~ Bixby”). In FIG. 27 (b), according to one embodiment, the electronic device (101) may confirm a standing state in which the electronic device (101) is mounted on a stand. In FIG. 27 (c), according to one embodiment, the electronic device (101) may operate in response mode based on the user’s gaze direction being directed toward the electronic device (101) while in a standing state. In FIG. 27 (d), the user may look in a different direction while the electronic device (101) is responding. According to one embodiment, the electronic device (101) may switch to listening mode after outputting all responses in response mode based on confirming the user’s gaze looking in a different direction. According to one embodiment, the electronic device (101) may stop outputting a response in response mode and switch to listening mode based on confirming that the user is looking in a different direction. In FIG. 27 (e), according to one embodiment, the electronic device (101) may maintain listening mode based on confirming that the user is not looking at the electronic device (101). In FIG. 27 (f), according to one embodiment, the electronic device (101) may disable the camera (220) and stop conversation mode based on the fact that voice data is not acquired for a specified period of time. In FIG. 27 (g), according to one embodiment, the electronic device (101) may reactivate conversation mode based on confirming a call to the electronic device (101) (e.g., “Hi~ Bixby”). In (h) of FIG. 27, according to one embodiment, when the conversation mode is reactivated, the electronic device (101) can use voice data stored before the conversation mode was stopped while performing the conversation mode.For example, the electronic device (101) can check response information generated using voice data stored before the conversation mode is stopped, and output a response (e.g., sound) based on the checked response information.
[0242] FIG. 28 is a flowchart of a method of operation of an electronic device according to one embodiment of the present disclosure. FIG. 28 may be described based on the embodiments of FIG. 1 to 27, the embodiments of FIG. 29, and embodiments described below. FIG. 29 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure.
[0243] Referring to FIG. 28, according to one embodiment, an electronic device (101) can check a security level per user and determine a level of response based on the security level.
[0244] At least some of the operations of FIG. 28 may be omitted. The order of the operations of FIG. 28 may be changed. Operations other than those of FIG. 28 may be performed before, during, or after the operations of FIG. 28.
[0245] Referring to FIG. 28, in operation 2801, according to one embodiment, an electronic device (101) can identify a user's biometric information. The biometric information may be information about the user's body identified in image data. For example, the biometric information may include information about the user's face. For example, the electronic device (101) may identify a face region in image data acquired through a camera (220) and obtain biometric information from the face region. For example, in FIG. 29, the electronic device (101) may identify feature points of the face in image data acquired through a camera (220) and obtain biometric information about the user's face based on the identified feature points.
[0246] In operation 2803, according to one embodiment, the electronic device (101) can determine the security level of a user based on biometric information. For example, the electronic device (101) can store information about the security level for each user in memory (130). The electronic device (101) can store matching information about the correspondence relationship between the user's biometric information and the security level in memory (130). For example, the electronic device (101) can determine the security level corresponding to the user by checking the security level that matches the biometric information confirmed in the image data obtained through the camera (220) in the matching information.
[0247] In operation 2805, according to one embodiment, the electronic device (101) may generate response information based on a security level. The security level may be a criterion for determining the level of the response. For example, the electronic device (101) may generate response information at a level corresponding to the user's security level in response mode. As the security level increases, the amount of information included in the response information may increase, or the amount of security information included in the response information may increase. The security information may be information set to be disclosed only to registered users. As the security level decreases, the amount of information included in the response information may decrease, or the amount of security information included in the response information may decrease. When the security level is the lowest level, security information may not be included in the response information. According to the embodiment of FIG. 28, the electronic device (101) may recognize the user's face, generate response information at different levels according to the security level for each user, and output a response (e.g., sound) based on the generated response information. According to one embodiment, the electronic device (101) may output a message (2910) (e.g., screen or sound) requesting registration based on the fact that the face of the user (2900) in FIG. 29 is not a registered face.
[0248] FIG. 30 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 30 may be described based on the embodiments of FIG. 1 to 29 and the embodiments described below.
[0249] In FIG. 30, according to one embodiment, an electronic device (101) can identify a primary user based on face area in image data acquired through a camera (220). For example, the electronic device (101) can identify face regions (e.g., 3010, 3020, 3030) corresponding to a plurality of users in the image data. The electronic device (101) can identify the location and area of the face regions. The electronic device (101) can set a primary user based on the location and area of the face regions. For example, the electronic device (101) can identify a face region (e.g., 3010) that is located in the middle of the image data and / or has the largest area, and set the user corresponding to the identified face region (e.g., 3010) as the primary user. The above-described embodiments may be applied to the primary user.
[0250] FIG. 31 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 31 may be described based on the embodiments of FIG. 1 to FIG. 30 and the embodiments described below.
[0251] According to one embodiment, the electronic device (101) can determine the direction in which a user's speech occurs based on voice data. In FIG. 31, the first user (3110) may be located on the left side of the electronic device (101), and the second user (3120) may be located on the right side of the electronic device (101). The electronic device (101) can determine the first speech of the first user (3110) and the second speech of the second user (3120) in the voice data, and determine the direction of the first speech and the direction of the second speech. According to one embodiment, the electronic device (101) can change the display of the screen based on the direction of the speech. For example, in FIG. 31, the electronic device (101) can display an icon in the direction of the speech. In FIG. 31 (a), the electronic device (101) can display a screen including an object (3101) indicating the execution of the group mode based on the entry into the group mode. In FIG. 31 (b), the electronic device (101) can determine the first direction of the first utterance based on confirming the first utterance of the first user (3110). The electronic device (101) can display an object (3102) in the first direction of the first utterance. In FIG. 31 (c), the electronic device (101) can determine the second direction of the second utterance based on confirming the second utterance of the second user (3120). The electronic device (101) can display an object (3103) in the second direction of the second utterance.
[0252] FIG. 32 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 32 may be described based on the embodiments of FIG. 1 to FIG. 31 and the embodiments described below.
[0253] The electronic device (101) of FIG. 32 may be a foldable device. For example, the electronic device (101) may include a first camera (3201) facing a first direction and a second camera (3202) facing a second direction opposite to the first direction. According to one embodiment, the electronic device (101) may, in a group mode where a plurality of users participate in a conversation, use the first camera (3201) to identify the first gaze direction of a first user (3210) and use the second camera (3202) to identify the second gaze direction of a second user (3220). According to one embodiment, the electronic device (101) may identify the first utterance of the first user (3210) in voice data. The electronic device (101) may select a response mode as an operation mode based on the fact that, while the first utterance of the first user (3210) is being confirmed, the first direction of gaze of the first user (3210) is a first direction toward the electronic device (101) (e.g., the screen of the display (250)). For example, the electronic device (101) may select a listening mode as an operation mode based on the fact that, while the first utterance of the first user (3210) is being confirmed, the first direction of gaze of the first user (3210) is not a first direction toward the electronic device (101) (e.g., the screen of the display (250)). For example, the electronic device (101) may ignore information regarding the second direction of gaze of the second user (3220) while the first utterance of the first user (3210) is being confirmed.
[0254] FIG. 33 is a drawing illustrating the operation of an electronic device according to one embodiment of the present disclosure. FIG. 33 may be described based on the embodiments of FIG. 1 to FIG. 32 and the embodiments described below.
[0255] In FIG. 33, according to one embodiment, the electronic device (101) may establish a communication connection with an external device (e.g., 3310, 3320, or 3330) (e.g., a wearable device) through a communication circuit (260). For example, the external device (e.g., 3310, 3320, or 3330) may include a smart necklace (3310) of FIG. 33 (a), a smart watch (3320) of FIG. 33 (b), or a smart ring (3330) of FIG. 33 (c), and there is no limitation on the type of external device (e.g., 3310, 3320, or 3330). An external device (e.g., 3310, 3320, or 3330) can measure the distance between the external device (e.g., 3310, 3320, or 3330) and the user and transmit information about the measured distance to the electronic device (101). The electronic device (101) can determine the distance between the external device (e.g., 3310, 3320, or 3330) and the user. According to one embodiment, the electronic device (101) can select a listening mode as an operating mode based on the distance between the external device (e.g., 3310, 3320, or 3330) and the user being less than a reference distance. For example, the electronic device (101) may receive additional voice input instead of outputting a response related to voice data received from the external device (e.g., 3310, 3320, or 3330) based on the distance between the external device (e.g., 3310, 3320, or 3330) and the user being less than a reference distance. According to one embodiment, the electronic device (101) may select a listening mode as an operating mode based on the first distance between the electronic device (101) and the user being greater than or equal to a reference distance and the second distance between the external device (e.g., 3310, 3320, or 3330) and the user being less than a reference distance.For example, the electronic device (101) may receive additional voice input instead of outputting a response related to voice data received from an external device (e.g., 3310, 3320, or 3330) based on the fact that a first distance between the electronic device (101) and the user is greater than or equal to a reference distance and a second distance between the external device (e.g., 3310, 3320, or 3330) and the user is less than a reference distance. According to one embodiment, the electronic device (101) may select a response mode as an operating mode based on the fact that the distance (e.g., second distance) between the external device (e.g., 3310, 3320, or 3330) and the user is greater than or equal to a reference distance. For example, the electronic device (101) may output a response related to voice data received from an external device (e.g., 3310, 3320, or 3330) based on the fact that the distance (e.g., a second distance) between the external device (e.g., 3310, 3320, or 3330) and the user is greater than or equal to a reference distance. According to one embodiment, the electronic device (101) may select a response mode as an operation mode based on the fact that the first distance between the electronic device (101) and the user is greater than or equal to a reference distance and the second distance between the external device (e.g., 3310, 3320, or 3330) and the user is greater than or equal to a reference distance. For example, the electronic device (101) can output a response related to voice data received from an external device (e.g., 3310, 3320, or 3330) based on the fact that a first distance between the electronic device (101) and the user is greater than or equal to a reference distance and a second distance between the external device (e.g., 3310, 3320, or 3330) and the user is greater than or equal to a reference distance.
[0256] In the embodiments described above, the response is explained as being output as sound, but this is exemplary, and the response of the electronic device (101) may be output in the form of text or image through the screen of the display (250).
[0257] Those skilled in the art will understand that the embodiments described herein may be applied interchangeably to the extent applicable. For example, those skilled in the art will understand that at least some operations of an embodiment described herein may be omitted, and at least some operations of the embodiments may be applied interchangeably.
[0258] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.
[0259] The present disclosure is not limited to the foregoing, and other unmentioned variations will be apparent to those skilled in the art from the present disclosure.
[0260] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.
[0261] According to one embodiment, the electronic device (101) may include a housing (200) forming the exterior of the electronic device (101), a display (250) disposed on a first surface of the housing (200), a camera (220) disposed in the direction facing the display (250), a microphone (210), at least one processor (120) including a processing circuit, and a memory (130) for storing instructions. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to operate in a conversational mode, in which a voice agent executed by the electronic device (101) interacts with voice input received through the microphone (210). When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to check whether at least one condition for outputting a first response associated with the first voice input is satisfied after receiving a first voice input. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to output the first response associated with the first voice input based on the fact that the at least one condition for outputting the first response is satisfied. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to receive a second voice input following the first voice input instead of outputting the first response based on the fact that the at least one condition for outputting the first response is not satisfied.The at least one condition for outputting the first response may include an operation of confirming that the user's gaze of the electronic device (101), obtained using the camera (220), moves to the screen of the display (250).
[0262] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to receive the second voice input following the first voice input instead of outputting the first response, based on the fulfillment of a second condition. The second condition may include an operation to confirm that the user's gaze, obtained using the camera (220), moves out of the screen of the display (250).
[0263] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to check whether at least one condition for outputting a second response related to the first voice input and the second voice input is satisfied after receiving the second voice input without outputting the first response. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to output the second response related to the first voice input and the second voice input based on the satisfaction of the at least one condition for outputting the second response. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to receive a third voice input following the second voice input instead of outputting the second response, based on the fact that the at least one condition for outputting the second response is not satisfied.
[0264] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to check whether the at least one condition for outputting the first response is satisfied based on confirming that no subsequent voice input is received for a specified period from the time the first voice input is received. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to output the first response associated with the first voice input based on the at least one condition for outputting the first response being satisfied.
[0265] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the user to continuously look toward the screen of the display (250) from before receiving the first voice input until after receiving the first voice input. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the user to continuously look toward the screen of the display (250), and the electronic device (101) may cause the electronic device to output the first response associated with the first voice input based on confirming that no subsequent voice input is received for a specified period from the time the first voice input is received.
[0266] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the camera (220) to be activated based at least on the activation of the conversation mode. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the camera (220) to be deactivated based on the deactivation of the conversation mode.
[0267] According to one embodiment, the at least one condition for outputting the first response may include an operation of verifying the designation assigned to the voice agent in the first voice input.
[0268] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to set a point corresponding to the user's eye in an image obtained using the camera (220) as a tracking point. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to identify the user's gaze by tracking the location of the tracking point corresponding to the user's eye. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to ignore other gazes based on identifying other gazes of a person other than the user.
[0269] According to one embodiment, the at least one condition for outputting the first response may include an operation to confirm that the first distance between the electronic device (101) and the user is greater than or equal to a reference distance.
[0270] According to one embodiment, the electronic device (101) may include at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to determine the positioning state of the electronic device (101) using the sensing value of the at least one sensor (231). The positioning state may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface, a standing state in which the back surface is placed at a certain angle to the floor surface, and a handheld state in which the electronic device (101) is held by a user. The at least one condition for outputting the first response may include an operation to check whether, in the floor state, the first voice input satisfies a first condition regarding whether it includes switching information that causes the conversation mode to switch operation modes. The at least one condition for outputting the first response may include an operation to check whether, in the standing state, the first condition or a second condition regarding the user's gaze direction is satisfied. The at least one condition for outputting the first response may include an operation to check whether, in the gripping state, the first condition, the second condition, or a third condition regarding the first distance between the electronic device (101) and the user is satisfied.
[0271] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to acquire image data necessary for determining the second condition by activating the camera (220) of the electronic device (101) based on confirming the standing state or the gripping state. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to deactivate the camera (220) based on a transition to the floor state, a failure to track the gaze direction, or the termination of the conversation mode.
[0272] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to perform an operation of measuring the first distance between the electronic device (101) and the user based on confirming the gripping state. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to stop measuring the first distance based on a transition to the floor state or the standing state, or the termination of the conversation mode.
[0273] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to enter the group mode based on confirming a second event that causes entry into a group mode in which a plurality of users participate in a conversation. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to confirm input for titles corresponding to the plurality of users so as to assign titles corresponding to the plurality of users respectively in the group mode. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to receive the second voice input following the first voice input instead of outputting the first response in the group mode, based on confirming at least one of the titles corresponding to the plurality of users in the first voice input.
[0274] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the user to confirm a gesture using the camera (220). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the first response associated with the first voice input to be output based on the fact that the gesture is a first gesture while the gaze toward the screen of the display (250) is not confirmed. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the second voice input following the first voice input to be received based on the fact that the gesture is a second gesture while the gaze toward the screen of the display (250) is not confirmed. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the conversation mode to end based on the fact that the gesture is a third gesture while the gaze toward the screen of the display (250) is not confirmed.
[0275] According to one embodiment, the electronic device (101) may include a foldable device. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the first user’s first gaze direction to be identified using the first camera (3201) of the electronic device (101) in a group mode in which multiple users participate in a conversation. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the second user’s second gaze direction to be identified using the second camera (3202) of the electronic device (101) in the group mode. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the first user’s first utterance to be identified in the voice input. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to output a response related to the first utterance based on the first gaze direction being directed toward the screen of the display (250) while the first utterance is being acknowledged. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to ignore information regarding the second gaze direction of the second user while the first utterance is being acknowledged.
[0276] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to establish a communication connection between the electronic device (101) and a wearable device through the communication circuit (260) of the electronic device (101). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to obtain voice data regarding the user's speech from the wearable device. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to determine a second distance between the wearable device and the user. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to output a response related to the voice data received from the wearable device based on the first distance being greater than or equal to the reference distance and the second distance being greater than or equal to the reference distance. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to receive additional voice input instead of outputting the response related to the voice data received from the wearable device based on the first distance being greater than or equal to the reference distance and the second distance being less than the reference distance.
[0277] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to set a point corresponding to the user's mouth in an image obtained using the camera (220) of the electronic device (101) as a tracking point. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to confirm the movement of the user's mouth by tracking the location of the tracking point corresponding to the user's mouth. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to ignore voice input that is confirmed while the movement of the user's mouth is not confirmed.
[0278] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the display (250) of the electronic device (101) to display a button for obtaining input. The at least one condition for outputting the first response may include an operation to confirm input through the button.
[0279] According to one embodiment, a method of operation of an electronic device (101) may include an operation of operating in a conversational mode in which a voice agent executed by the electronic device (101) interacts with a voice input received through a microphone (210). The method may include an operation of checking whether at least one condition for outputting a first response related to the first voice input is satisfied after receiving a first voice input. The method may include an operation of outputting the first response related to the first voice input based on the satisfaction of the at least one condition for outputting the first response. The method may include an operation of receiving a second voice input following the first voice input instead of outputting the first response based on the failure to satisfy the at least one condition for outputting the first response. The at least one condition for outputting the first response may include an operation of checking that the gaze of the user of the electronic device (101), obtained using the camera (220), moves to the screen of the display (250).
[0280] According to one embodiment, the method may include an operation of receiving the second voice input following the first voice input instead of outputting the first response, based on the second condition being satisfied. The second condition may include an operation of confirming that the user's gaze, obtained using the camera (220), moves outside the screen of the display (250).
[0281] According to one embodiment, the method may include an operation of checking whether at least one condition for outputting a second response related to the first voice input and the second voice input is satisfied after receiving the second voice input without outputting the first response. The method may include an operation of outputting the second response related to the first voice input and the second voice input based on the satisfaction of the at least one condition for outputting the second response. The method may include an operation of receiving a third voice input following the second voice input instead of outputting the second response based on the fact that the at least one condition for outputting the second response is not satisfied.
[0282] According to one embodiment, the method may include an operation of checking whether the at least one condition for outputting the first response is satisfied based on confirming that a subsequent voice input is not received for a specified period from the time when the first voice input is received. The method may include an operation of outputting the first response associated with the first voice input based on the satisfaction of the at least one condition for outputting the first response.
[0283] According to one embodiment, the method may include an operation of checking the continuous gaze of the user toward the screen of the display (250) from before receiving the first voice input until after receiving the first voice input. The method may include an operation of outputting the first response related to the first voice input based on confirming that no subsequent voice input is received for a specified period from the time the first voice input is received, while checking the continuous gaze of the user toward the screen of the display (250).
[0284] According to one embodiment, the method may include an operation to activate the camera (220) based at least on the conversation mode being activated. The method may include an operation to deactivate the camera (220) based on the conversation mode being deactivated.
[0285] According to one embodiment, in the method, the at least one condition for outputting the first response may include an operation of verifying the designation assigned to the voice agent in the first voice input.
[0286] According to one embodiment, the method may include an operation of setting a point corresponding to the user's eye in an image obtained using the camera (220) as a tracking point. The method may include an operation of confirming the user's gaze by tracking the location of the tracking point corresponding to the user's eye. The method may include an operation of ignoring other gazes based on confirming other gazes of a person other than the user.
[0287] According to one embodiment, in the method, the at least one condition for outputting the first response may include an operation of confirming that the first distance between the electronic device (101) and the user is greater than or equal to a reference distance.
[0288] According to one embodiment, the method may include an operation of determining the positioning state of the electronic device (101) using a sensing value of at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101). The positioning state may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface, a standing state in which the back surface is placed at a certain angle to the floor surface, and a handheld state in which the electronic device (101) is held by a user. The at least one condition for outputting the first response may include an operation of determining whether, in the floor state, the first voice input satisfies a first condition regarding whether it contains switching information that causes the operation mode of the conversation mode to be switched. The at least one condition for outputting the first response may include an operation of determining whether, in the standing state, the first condition or a second condition regarding the direction of the user's gaze is satisfied. The at least one condition for outputting the first response may include an operation to check whether the first condition, the second condition, or the third condition regarding the first distance between the electronic device (101) and the user is satisfied in the gripping state.
[0289] According to one embodiment, the method may include an operation of acquiring image data necessary for determining the second condition by activating the camera (220) of the electronic device (101) based on confirming the standing state or the gripping state. The method may include an operation of deactivating the camera (220) based on a transition to the floor state, failure of tracking the gaze direction, or termination of the conversation mode.
[0290] According to one embodiment, the method may include an operation to measure the first distance between the electronic device (101) and the user based on confirming the gripping state. The method may include an operation to stop measuring the first distance based on a transition to the floor state or the standing state, or the termination of the conversation mode.
[0291] According to one embodiment, the method may include an operation of entering a group mode based on confirming a second event that causes entry into a group mode in which a plurality of users participate in a conversation. The method may include an operation of confirming input for titles corresponding to a plurality of users so as to assign titles corresponding to each (respectively) the plurality of users in the group mode. The method may include an operation of receiving a second voice input following the first voice input instead of outputting the first response, based on confirming at least one of the titles corresponding to the plurality of users in the first voice input in the group mode.
[0292] According to one embodiment, the method may include an operation of confirming the user's gesture using the camera (220). The method may include an operation of outputting the first response related to the first voice input based on the fact that the gesture is a first gesture while the gaze toward the screen of the display (250) is not confirmed. The method may include an operation of receiving the second voice input following the first voice input instead of outputting the first response based on the fact that the gesture is a second gesture while the gaze toward the screen of the display (250) is not confirmed. The method may include an operation of terminating the conversation mode based on the fact that the gesture is a third gesture while the gaze toward the screen of the display (250) is not confirmed.
[0293] According to one embodiment, in the method, the electronic device (101) may include a foldable device. The method may include an operation of confirming a first gaze direction of a first user using a first camera (3201) of the electronic device (101) in a group mode in which a plurality of users participate in a conversation. The method may include an operation of confirming a second gaze direction of a second user using a second camera (3202) of the electronic device (101) in the group mode. The method may include an operation of confirming a first utterance of the first user in the voice input. The method may include an operation of outputting a response related to the first utterance based on the first gaze direction being directed toward the screen of the display (250) while the first utterance is being confirmed. The method may include an operation of ignoring information regarding the second gaze direction of the second user while the first utterance is being confirmed.
[0294] According to one embodiment, the method may include an operation of establishing a communication connection between the electronic device (101) and the wearable device through a communication circuit (260) of the electronic device (101). The method may include an operation of obtaining voice data regarding the user's speech from the wearable device. The method may include an operation of determining a second distance between the wearable device and the user. The method may include an operation of outputting a response related to the voice data received from the wearable device based on the fact that the first distance is greater than or equal to the reference distance and the second distance is greater than or equal to the reference distance. The method may include an operation of receiving additional voice input instead of outputting the response related to the voice data received from the wearable device based on the fact that the first distance is greater than or equal to the reference distance and the second distance is less than the reference distance.
[0295] According to one embodiment, the method may include an operation of setting a point corresponding to the user's mouth in an image obtained using the camera (220) of the electronic device (101) as a tracking point. The method may include an operation of confirming the movement of the user's mouth by tracking the location of the tracking point corresponding to the user's mouth. The method may include an operation of ignoring voice input that is detected while the movement of the user's mouth is not confirmed.
[0296] According to one embodiment, the method may include an operation of controlling the display (250) of the electronic device (101) to display a button for obtaining input. The at least one condition for outputting the first response may include an operation of confirming input through the button.
[0297] According to one embodiment, in a non-transitory computer-readable recording medium for storing instructions, the instructions may cause the electronic device (101) to perform at least one operation when executed individually or collectively by at least one processor of the electronic device (101). The at least one operation may include an operation of operating in a conversational mode, in which a voice agent executed by the electronic device (101) interacts with a voice input received through a microphone (210). The at least one operation may include an operation of checking whether at least one condition for outputting a first response associated with the first voice input is satisfied after receiving the first voice input. The at least one operation may include an operation of outputting the first response associated with the first voice input based on the satisfaction of the at least one condition for outputting the first response. The above at least one operation may include receiving a second voice input following the first voice input instead of outputting the first response, based on the fact that the above at least one condition for outputting the first response is not satisfied. The above at least one condition for outputting the first response may include confirming that the user's gaze of the electronic device (101), obtained using the camera (220), moves to the screen of the display (250).
[0298] According to one embodiment, in the recording medium, the at least one operation may include receiving the second voice input following the first voice input instead of outputting the first response, based on the second condition being satisfied. The second condition may include confirming that the user's gaze, obtained using the camera (220), moves outside the screen of the display (250).
[0299] According to one embodiment, in the recording medium, the at least one operation may include an operation of checking whether at least one condition for outputting a second response related to the first voice input and the second voice input is satisfied after receiving the second voice input without outputting the first response. The at least one operation may include an operation of outputting the second response related to the first voice input and the second voice input based on the satisfaction of the at least one condition for outputting the second response. The at least one operation may include an operation of receiving a third voice input following the second voice input instead of outputting the second response based on the fact that the at least one condition for outputting the second response is not satisfied.
[0300] According to one embodiment, in the recording medium, the at least one operation may include an operation of checking whether the at least one condition for outputting the first response is satisfied based on confirming that a subsequent voice input is not received for a specified period from the time when the first voice input is received. The at least one operation may include an operation of outputting the first response associated with the first voice input based on the satisfaction of the at least one condition for outputting the first response.
[0301] According to one embodiment, in the recording medium, the at least one operation may include an operation of checking the continuous gaze of the user toward the screen of the display (250) from before receiving the first voice input until after receiving the first voice input. The at least one operation may include an operation of outputting the first response related to the first voice input based on confirming that a subsequent voice input is not received for a specified period from the time the first voice input is received while checking the continuous gaze of the user toward the screen of the display (250).
[0302] According to one embodiment, in the recording medium, the at least one operation may include an operation to activate the camera (220) based at least on the activation of the conversation mode. The at least one operation may include an operation to deactivate the camera (220) based on the deactivation of the conversation mode.
[0303] According to one embodiment, in the recording medium, the at least one condition for outputting the first response may include an operation of verifying the designation assigned to the voice agent in the first voice input.
[0304] According to one embodiment, in the recording medium, the at least one operation may include setting a point corresponding to the user's eye in an image obtained using the camera (220) as a tracking point. The at least one operation may include confirming the user's gaze by tracking the location of the tracking point corresponding to the user's eye. The at least one operation may include ignoring the other gaze based on confirming the other gaze of a person other than the user.
[0305] According to one embodiment, in the recording medium, the at least one condition for outputting the first response may include an operation of confirming that the first distance between the electronic device (101) and the user is greater than or equal to a reference distance.
[0306] According to one embodiment, in the recording medium, the at least one operation may include an operation of determining the positioning state of the electronic device (101) using a sensing value of at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101). The positioning state may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface, a standing state in which the back surface is placed at a certain angle to the floor surface, and a handheld state in which the electronic device (101) is held by a user. The at least one condition for outputting the first response may include an operation of determining whether, in the floor state, the first voice input satisfies a first condition regarding whether it includes switching information that causes the operation mode of the conversation mode to be switched. The at least one condition for outputting the first response may include an operation to check whether the first condition or the second condition regarding the user's gaze direction is satisfied in the standing state. The at least one condition for outputting the first response may include an operation to check whether the first condition, the second condition, or the third condition regarding the first distance between the electronic device (101) and the user is satisfied in the gripping state.
[0307] According to one embodiment, in the recording medium, the at least one operation may include an operation of acquiring image data necessary for determining the second condition by activating the camera (220) of the electronic device (101) based on confirming the standing state or the gripping state. The at least one operation may include an operation of deactivating the camera (220) based on a transition to the floor state, failure of tracking the gaze direction, or termination of the conversation mode.
[0308] According to one embodiment, in the recording medium, the at least one operation may include an operation to measure the first distance between the electronic device (101) and the user based on confirming the gripping state. The at least one operation may include an operation to stop measuring the first distance based on a transition to the floor state or the standing state, or the termination of the conversation mode.
[0309] According to one embodiment, in the recording medium, the at least one operation may include an operation of entering the group mode based on confirming a second event that causes entry into the group mode in which a plurality of users participate in a conversation. The at least one operation may include an operation of confirming input for titles corresponding to the plurality of users so as to assign titles corresponding to the plurality of users respectively in the group mode. The at least one operation may include an operation of receiving the second voice input following the first voice input instead of outputting the first response based on confirming at least one of the titles corresponding to the plurality of users in the first voice input in the group mode.
[0310] According to one embodiment, in the recording medium, the at least one operation may include an operation of confirming the user's gesture using the camera (220). The at least one operation may include an operation of outputting the first response related to the first voice input based on the fact that the gesture is a first gesture while the gaze toward the screen of the display (250) is not confirmed. The at least one operation may include an operation of receiving the second voice input following the first voice input instead of outputting the first response based on the fact that the gesture is a second gesture while the gaze toward the screen of the display (250) is not confirmed. The at least one operation may include an operation of terminating the conversation mode based on the fact that the gesture is a third gesture while the gaze toward the screen of the display (250) is not confirmed.
[0311] According to one embodiment, in the recording medium, the electronic device (101) may include a foldable device. The at least one operation may include an operation of confirming the first gaze direction of a first user using the first camera (3201) of the electronic device (101) in a group mode in which a plurality of users participate in a conversation. The at least one operation may include an operation of confirming the second gaze direction of a second user using the second camera (3202) of the electronic device (101) in the group mode. The at least one operation may include an operation of confirming the first utterance of the first user in the voice input. The at least one operation may include an operation of outputting a response related to the first utterance based on the first gaze direction being directed toward the screen of the display (250) while the first utterance is being confirmed. The at least one operation may include an operation of ignoring information regarding the second gaze direction of the second user while the first utterance is being confirmed.
[0312] According to one embodiment, in the recording medium, the at least one operation may include an operation of establishing a communication connection between the electronic device (101) and the wearable device through a communication circuit (260) of the electronic device (101). The at least one operation may include an operation of obtaining voice data for the user's speech from the wearable device. The at least one operation may include an operation of determining a second distance between the wearable device and the user. The at least one operation may include an operation of outputting a response related to the voice data received from the wearable device based on the fact that the first distance is greater than or equal to the reference distance and the second distance is greater than or equal to the reference distance. The at least one operation may include an operation of receiving additional voice input instead of outputting the response related to the voice data received from the wearable device based on the fact that the first distance is greater than or equal to the reference distance and the second distance is less than the reference distance.
[0313] According to one embodiment, in the recording medium, the at least one operation may include setting a point corresponding to the user's mouth in an image obtained using the camera (220) of the electronic device (101) as a tracking point. The at least one operation may include confirming the movement of the user's mouth by tracking the location of the tracking point corresponding to the user's mouth. The at least one operation may include ignoring voice input that is detected while the movement of the user's mouth is not confirmed.
[0314] According to one embodiment, in the recording medium, the at least one operation may include an operation of controlling the display (250) of the electronic device (101) to display a button for acquiring input. The at least one condition for outputting the first response may include an operation of confirming input through the button. According to one embodiment, the electronic device (101) may include a microphone (210) configured to acquire voice data, at least one processor (120) including a processing circuit, and a memory (130) for storing instructions. When the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) may be caused to enter the conversation mode based on confirming a first event that causes entry into the conversation mode with an artificial intelligence model (e.g., a voice agent). When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to determine the operating mode among a listening mode or a response mode based on conditions for determining the operating mode of the conversation mode. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to perform the operation of the listening mode or the operation of the response mode based on determining the listening mode or the response mode as the operating mode. The operation of the listening mode may include the operation of acquiring the voice data. The operation of the response mode may include the operation of acquiring the voice data, the operation of verifying response information generated based on the voice data accumulated during the conversation mode, and the operation of outputting a response based on the response information.
[0315] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to check whether the voice data includes switching information that causes the operation mode of the conversation mode to be switched. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the response mode as the operation mode based on checking the designation assigned to the artificial intelligence model (e.g., voice agent) in the voice data. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the listening mode as the operation mode based on checking the designation assigned to the user in the voice data.
[0316] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the user’s gaze direction to be identified using the camera (220) of the electronic device (101). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the response mode to be selected as the operation mode based on identifying the user’s gaze in a first direction toward the electronic device (101) (e.g., the screen of the display (250)). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the listening mode to be selected as the operation mode based on failing to identify the gaze in the first direction.
[0317] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to set a point corresponding to the user's eye in an image obtained using the camera (220) as a tracking point. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to identify the user's gaze by tracking the location of the tracking point corresponding to the user's eye. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to ignore other gazes based on identifying other gazes of a person other than the user.
[0318] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to determine a first distance between the electronic device (101) and the user. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the listening mode to be selected as the operating mode based on the first distance being less than a reference distance. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the response mode to be selected as the operating mode based on the first distance being greater than or equal to the reference distance.
[0319] According to one embodiment, the electronic device (101) may include at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to determine the positioning state of the electronic device (101) using the sensing value of the at least one sensor (231). The positioning state may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface. The positioning state may include a standing state in which the back surface is placed at a certain angle to the floor surface. The positioning state may include a handheld state in which the electronic device (101) is held by a user. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to determine the listening mode or the response mode as the operating mode based on a first condition regarding whether the voice data includes switching information that causes the conversation mode to switch the operating mode. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to determine the listening mode or the response mode as the operating mode based on the first condition and a second condition regarding the user's gaze direction in the standing state.When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to determine the listening mode or the response mode as the operating mode based on the first condition, the second condition, and the third condition regarding the first distance between the electronic device (101) and the user in the gripping state.
[0320] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to acquire image data necessary for determining the second condition by activating the camera (220) of the electronic device (101) based on confirming the standing state or the gripping state. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to deactivate the camera (220) based on a transition to the floor state, a failure of tracking the gaze direction, or the termination of the conversation mode.
[0321] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to perform an operation of measuring the first distance between the electronic device (101) and the user based on confirming the gripping state. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to stop measuring the first distance based on a transition to the floor state or the standing state, or the termination of the conversation mode.
[0322] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to enter the group mode based on confirming a second event that causes entry into a group mode in which a plurality of users participate in a conversation. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to confirm input for the titles corresponding to the plurality of users so as to assign titles corresponding to the plurality of users respectively in the group mode. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the response mode as the operation mode based on confirming the title assigned to the artificial intelligence model (e.g., voice agent) in the voice data in the group mode. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the listening mode as the operating mode based on identifying at least one of the names corresponding to the plurality of users in the voice data in the group mode.
[0323] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the user's gesture to be identified using the camera (220). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the response mode to be selected as the operation mode based on the fact that the gesture is the first gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not identified. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the listening mode as the operating mode based on the fact that the gesture is a second gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to terminate the conversation mode based on the fact that the gesture is a third gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed.
[0324] According to one embodiment, the electronic device (101) may include a foldable device. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the first user’s first gaze direction to be identified using the first camera (3201) of the electronic device (101) in a group mode in which multiple users participate in a conversation. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the second user’s second gaze direction to be identified using the second camera (3202) of the electronic device (101) in the group mode. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the first user’s first utterance to be identified in the voice data. When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the response mode as the operation mode based on the fact that, while the first utterance is being acknowledged, the first line of sight is the first direction toward the electronic device (101) (e.g., the screen of the display (250)). When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the listening mode as the operation mode based on the fact that, while the first utterance is being acknowledged, the first line of sight is not the first direction.When the above instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to ignore information regarding the second user's second line of sight direction while the first utterance is being confirmed.
[0325] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to establish a communication connection between the electronic device (101) and a wearable device through the communication circuit (260) of the electronic device (101). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to obtain voice data regarding the user's speech from the wearable device. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the electronic device (101) to determine a second distance between the wearable device and the user. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to select the listening mode as the operating mode based on the first distance being greater than or equal to the reference distance and the second distance being less than the reference distance. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to select the response mode as the operating mode based on the first distance being greater than or equal to the reference distance and the second distance being greater than or equal to the reference distance.
[0326] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the listening mode as the operating mode based on identifying a filler voice that causes the listening mode in the voice data.
[0327] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to set a point corresponding to the user's mouth in an image obtained using the camera (220) of the electronic device (101) as a tracking point. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to confirm the movement of the user's mouth by tracking the location of the tracking point corresponding to the user's mouth. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to ignore voice data that is confirmed while the movement of the user's mouth is not confirmed.
[0328] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the display (250) of the electronic device (101) to display a button for acquiring input. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the response mode to be selected as the operation mode based on confirming input through the button. When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may cause the operation mode to be selected based on other conditions based on the fact that the input through the button is not confirmed.
[0329] According to one embodiment, when the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to identify biometric information about the user's face in an image obtained using the camera (220). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to identify a security level corresponding to the user by identifying a security level that matches the biometric information in the matching information stored in the memory (130). When the instructions are executed individually or collectively by at least one processor (120), the electronic device (101) may be caused to select the response mode as the operation mode based on identifying the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)). When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to generate the response information at a level corresponding to the user's security level in the response mode. When the above instructions are executed individually or collectively by at least one processor (120), they may cause the electronic device (101) to output the response based on the response information.
[0330] According to one embodiment, a method of operation of an electronic device (101) may include an operation of entering a conversation mode based on confirming a first event that causes entry into a conversation mode with an artificial intelligence model (e.g., a voice agent). The method may include an operation of determining the operation mode among a listening mode or a response mode based on a condition for determining the operation mode of the conversation mode. The method may include an operation of performing the operation of the listening mode or the operation of the response mode based on determining the listening mode or the response mode as the operation mode. The operation of the listening mode may include an operation of acquiring voice data. The operation of the response mode may include an operation of acquiring the voice data, an operation of confirming response information generated based on the voice data accumulated during the conversation mode, and an operation of outputting a response based on the response information.
[0331] According to one embodiment, the method may include an action of checking whether the voice data includes switching information that causes the operation mode of the conversation mode to be switched. The method may include an action of selecting the response mode as the operation mode based on checking the designation assigned to the artificial intelligence model (e.g., voice agent) in the voice data. The method may include an action of selecting the listening mode as the operation mode based on checking the designation assigned to the user in the voice data.
[0332] According to one embodiment, the method may include an operation of checking the direction of the user's gaze using a camera (220) of the electronic device (101). The method may include an operation of selecting the response mode as the operation mode based on checking the user's gaze in a first direction toward the electronic device (101) (e.g., the screen of the display (250)). The method may include an operation of selecting the listening mode as the operation mode based on failing to check the gaze in the first direction.
[0333] According to one embodiment, the method may include an operation of setting a point corresponding to the user's eye in an image obtained using the camera (220) as a tracking point. The method may include an operation of confirming the user's gaze by tracking the location of the tracking point corresponding to the user's eye. The method may include an operation of ignoring other gazes based on confirming other gazes of a person other than the user.
[0334] According to one embodiment, the method may include an operation of checking a first distance between the electronic device (101) and a user. The method may include an operation of selecting the listening mode as the operating mode based on the first distance being less than a reference distance. The method may include an operation of selecting the response mode as the operating mode based on the first distance being greater than or equal to the reference distance.
[0335] According to one embodiment, in the method, the electronic device (101) may include at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101). The method may include an operation of determining the positioning state of the electronic device (101) using the sensing value of the at least one sensor (231). The positioning state may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface. The positioning state may include a standing state in which the back surface is placed at a certain angle to the floor surface. The positioning state may include a handheld state in which the electronic device (101) is held by a user. The method may include an operation of determining the listening mode or the response mode as the operating mode based on a first condition regarding whether the voice data includes switching information that causes the voice data to switch the operating mode of the conversation mode. The above method may include an operation of determining the listening mode or the response mode as the operation mode based on the first condition and the second condition regarding the direction of the user's gaze in the standing state. The above method may include an operation of determining the listening mode or the response mode as the operation mode based on the first condition, the second condition, and the third condition regarding the first distance between the electronic device (101) and the user in the holding state.
[0336] According to one embodiment, the method may include an operation of acquiring image data necessary for determining the second condition by activating the camera (220) of the electronic device (101) based on confirming the standing state or the gripping state. The method may include an operation of deactivating the camera (220) based on a transition to the floor state, failure of tracking the gaze direction, or termination of the conversation mode.
[0337] According to one embodiment, the method may include an operation to measure the first distance between the electronic device (101) and the user based on confirming the gripping state. The method may include an operation to stop measuring the first distance based on a transition to the floor state or the standing state, or the termination of the conversation mode.
[0338] According to one embodiment, the method may include an operation of entering a group mode based on confirming a second event that causes entry into a group mode in which a plurality of users participate in a conversation. The method may include an operation of confirming input for titles corresponding to a plurality of users so as to assign titles corresponding to each (respectively) the plurality of users in the group mode. The method may include an operation of selecting the response mode as the operation mode based on confirming the title assigned to the artificial intelligence model (e.g., voice agent) in the voice data in the group mode. The method may include an operation of selecting the listening mode as the operation mode based on confirming at least one of the titles corresponding to the plurality of users in the voice data in the group mode.
[0339] According to one embodiment, the method may include an action of confirming the user's gesture using the camera (220). The method may include an action of selecting the response mode as the action mode based on the fact that the gesture is a first gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed. The method may include an action of selecting the listening mode as the action mode based on the fact that the gesture is a second gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed. The method may include an action of terminating the conversation mode based on the fact that the gesture is a third gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed.
[0340] According to one embodiment, the electronic device (101) may include a foldable device. The method may include an operation of confirming a first user's first gaze direction using a first camera (3201) of the electronic device (101) in a group mode in which a plurality of users participate in a conversation. The method may include an operation of confirming a second user's second gaze direction using a second camera (3202) of the electronic device (101) in the group mode. The method may include an operation of confirming a first utterance of the first user in the voice data. The method may include an operation of selecting the response mode as the operation mode based on the fact that, while the first utterance is being confirmed, the first gaze direction is the first direction toward the electronic device (101) (e.g., the screen of the display (250)). The method may include an operation of selecting the listening mode as the operation mode based on the fact that, while the first utterance is being confirmed, the first gaze direction is not the first direction. The above method may include an operation of ignoring information regarding the second user's second gaze direction while the first utterance is being confirmed.
[0341] According to one embodiment, the method may include an operation of establishing a communication connection between the electronic device (101) and the wearable device through a communication circuit (260) of the electronic device (101). The method may include an operation of obtaining voice data regarding the user's speech from the wearable device. The method may include an operation of checking a second distance between the wearable device and the user. The method may include an operation of selecting the listening mode as the operating mode based on the fact that the first distance is greater than or equal to the reference distance and the second distance is less than the reference distance. The method may include an operation of selecting the response mode as the operating mode based on the fact that the first distance is greater than or equal to the reference distance and the second distance is greater than or equal to the reference distance.
[0342] According to one embodiment, the method may include an operation of selecting the listening mode as the operation mode based on identifying a filler voice that causes the listening mode in the voice data.
[0343] According to one embodiment, the method may include an operation of setting a point corresponding to the user's mouth in an image obtained using a camera (220) of the electronic device (101) as a tracking point. The method may include an operation of confirming the movement of the user's mouth by tracking the location of the tracking point corresponding to the user's mouth. The method may include an operation of ignoring voice data that is confirmed while the movement of the user's mouth is not confirmed.
[0344] According to one embodiment, the method may include an operation of controlling a display (250) of the electronic device (101) to display a button for acquiring input. The method may include an operation of selecting the response mode as the operation mode based on confirming the input through the button. The method may include an operation of selecting the operation mode based on other conditions based on the fact that the input through the button is not confirmed.
[0345] According to one embodiment, the method may include an operation of verifying biometric information of the user's face in an image obtained using the camera (220). The method may include an operation of verifying a security level corresponding to the user by verifying a security level that matches the biometric information in the matching information stored in the memory (130). The method may include an operation of selecting the response mode as the operation mode based on verifying the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)). The method may include an operation of generating response information at a level corresponding to the user's security level in the response mode. The method may include an operation of outputting the response based on the response information.
[0346] According to one embodiment, in a non-transitory computer-readable recording medium for storing instructions, the instructions may cause the electronic device (101) to perform at least one operation when executed individually or collectively by at least one processor of the electronic device (101). The at least one operation may include an operation of entering a conversation mode based on confirming a first event that causes entry into a conversation mode with an artificial intelligence model (e.g., a voice agent). The at least one operation may include an operation of determining the operation mode among a listening mode or a response mode based on a condition for determining the operation mode of the conversation mode. The at least one operation may include an operation of performing the operation of the listening mode or the operation of the response mode based on determining the listening mode or the response mode as the operation mode. The operation of the listening mode may include an operation of acquiring voice data. The operation of the above response mode may include the operation of acquiring the voice data, the operation of verifying response information generated based on the voice data accumulated during the conversation mode, and the operation of outputting a response based on the response information.
[0347] According to one embodiment, in the recording medium, the at least one operation may include an operation of checking whether the voice data includes switching information that causes the conversation mode to switch the operation mode. The at least one operation may include an operation of selecting the response mode as the operation mode based on checking the designation assigned to the artificial intelligence model (e.g., voice agent) in the voice data. The at least one operation may include an operation of selecting the listening mode as the operation mode based on checking the designation assigned to the user in the voice data.
[0348] According to one embodiment, in the recording medium, the at least one operation may include an operation of checking the direction of the user's gaze using a camera (220) of the electronic device (101). The at least one operation may include an operation of selecting the response mode as the operation mode based on checking the user's gaze in a first direction toward the electronic device (101) (e.g., the screen of the display (250)). The at least one operation may include an operation of selecting the listening mode as the operation mode based on failing to check the gaze in the first direction.
[0349] According to one embodiment, in the recording medium, the at least one operation may include setting a point corresponding to the user's eye in an image obtained using the camera (220) as a tracking point. The at least one operation may include confirming the user's gaze by tracking the location of the tracking point corresponding to the user's eye. The at least one operation may include ignoring the other gaze based on confirming the other gaze of a person other than the user.
[0350] According to one embodiment, in the recording medium, the at least one operation may include an operation of checking a first distance between the electronic device (101) and a user. The at least one operation may include an operation of selecting the listening mode as the operation mode based on the first distance being less than a reference distance. The at least one operation may include an operation of selecting the response mode as the operation mode based on the first distance being greater than or equal to the reference distance.
[0351] According to one embodiment, in the recording medium, the electronic device (101) may include at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101). The at least one operation may include an operation of determining the positioning state of the electronic device (101) using the sensing value of the at least one sensor (231). The positioning state may include a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface. The positioning state may include a standing state in which the back surface is placed at a certain angle with the floor surface. The positioning state may include a handheld state in which the electronic device (101) is held by a user. The at least one operation may include, in the floor state, an operation of determining the listening mode or the response mode as the operation mode based on a first condition regarding whether the voice data includes switching information that causes the conversation mode to switch the operation mode. The at least one operation may include, in the standing state, an operation of determining the listening mode or the response mode as the operation mode based on the first condition and a second condition regarding the direction of the user's gaze. The at least one operation may include, in the gripping state, an operation of determining the listening mode or the response mode as the operation mode based on the first condition, the second condition, and a third condition regarding the first distance between the electronic device (101) and the user.
[0352] According to one embodiment, in the recording medium, the at least one operation may include an operation of acquiring image data necessary for determining the second condition by activating the camera (220) of the electronic device (101) based on confirming the standing state or the gripping state. The at least one operation may include an operation of deactivating the camera (220) based on a transition to the floor state, failure of tracking the gaze direction, or termination of the conversation mode.
[0353] According to one embodiment, in the recording medium, the at least one operation may include an operation to measure the first distance between the electronic device (101) and the user based on confirming the gripping state. The at least one operation may include an operation to stop measuring the first distance based on a transition to the floor state or the standing state, or the termination of the conversation mode.
[0354] According to one embodiment, in the recording medium, the at least one operation may include an operation of entering the group mode based on confirming a second event that causes entry into a group mode in which a plurality of users participate in a conversation. The at least one operation may include an operation of confirming input for titles corresponding to the plurality of users so as to assign titles corresponding to the plurality of users respectively in the group mode. The at least one operation may include an operation of selecting the response mode as the operation mode based on confirming the title assigned to the artificial intelligence model (e.g., voice agent) in the voice data in the group mode. The at least one operation may include an operation of selecting the listening mode as the operation mode based on confirming at least one of the titles corresponding to the plurality of users in the voice data in the group mode.
[0355] According to one embodiment, in the recording medium, the at least one operation may include an operation of confirming the user's gesture using the camera (220). The at least one operation may include an operation of selecting the response mode as the operation mode based on the fact that the gesture is a first gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed. The at least one operation may include an operation of selecting the listening mode as the operation mode based on the fact that the gesture is a second gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed. The at least one operation may include an operation of ending the conversation mode based on the fact that the gesture is a third gesture while the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)) is not confirmed.
[0356] According to one embodiment, in the recording medium, the electronic device (101) may include a foldable device. The at least one operation may include an operation of confirming the first gaze direction of a first user using the first camera (3201) of the electronic device (101) in a group mode in which a plurality of users participate in a conversation. The at least one operation may include an operation of confirming the second gaze direction of a second user using the second camera (3202) of the electronic device (101) in the group mode. The at least one operation may include an operation of confirming the first utterance of the first user in the voice data. The at least one operation may include an operation of selecting the response mode as the operation mode based on the fact that, while the first utterance is being confirmed, the first gaze direction is the first direction toward the electronic device (101) (e.g., the screen of the display (250)). The at least one operation may include an operation of selecting the listening mode as the operation mode based on the fact that the first gaze direction is not the first direction while the first utterance is being confirmed. The at least one operation may include an operation of ignoring information regarding the second user's second gaze direction while the first utterance is being confirmed.
[0357] According to one embodiment, in the recording medium, the at least one operation may include an operation of establishing a communication connection between the electronic device (101) and the wearable device through a communication circuit (260) of the electronic device (101). The at least one operation may include an operation of obtaining voice data for the user's speech from the wearable device. The at least one operation may include an operation of checking a second distance between the wearable device and the user. The at least one operation may include an operation of selecting the listening mode as the operation mode based on the fact that the first distance is greater than or equal to the reference distance and the second distance is less than the reference distance. The at least one operation may include an operation of selecting the response mode as the operation mode based on the fact that the first distance is greater than or equal to the reference distance and the second distance is greater than or equal to the reference distance.
[0358] According to one embodiment, in the recording medium, the at least one operation may include an operation of selecting the listening mode as the operation mode based on identifying a filler voice that causes the listening mode in the voice data.
[0359] According to one embodiment, in the recording medium, the at least one operation may include setting a point corresponding to the user's mouth in an image obtained using the camera (220) of the electronic device (101) as a tracking point. The at least one operation may include confirming the movement of the user's mouth by tracking the location of the tracking point corresponding to the user's mouth. The at least one operation may include ignoring voice data confirmed while the movement of the user's mouth is not confirmed.
[0360] According to one embodiment, in the recording medium, the at least one operation may include an operation of controlling a display (250) of the electronic device (101) to display a button for acquiring input. The at least one operation may include an operation of selecting the response mode as the operation mode based on confirming the input through the button. The at least one operation may include an operation of selecting the operation mode based on other conditions based on the fact that the input through the button is not confirmed.
[0361] According to one embodiment, in the recording medium, the at least one operation may include an operation of verifying biometric information of the user's face in an image obtained using the camera (220). The at least one operation may include an operation of verifying a security level corresponding to the user by verifying a security level matching the biometric information in the matching information stored in the memory (130). The at least one operation may include an operation of selecting the response mode as the operation mode based on verifying the gaze in the first direction toward the electronic device (101) (e.g., the screen of the display (250)). The at least one operation may include an operation of generating response information at a level corresponding to the user's security level in the response mode. The at least one operation may include an operation of outputting the response based on the response information.
[0362] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0363] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0364] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0365] Various embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored on a storage medium readable by a machine (e.g., an electronic device). For example, a processor (e.g., a controller) of the machine may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0366] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0367] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0368] It will be understood that various embodiments of the present disclosure according to the claims and description of this specification may be implemented in the form of hardware, software, or a combination of hardware and software.
[0369] Such software may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores one or more computer programs (software modules), and the one or more computer programs include computer-executable instructions that cause the electronic device to perform the method of the present disclosure when executed individually or collectively by one or more processors of the electronic device.
[0370] Such software may be stored on a volatile or non-volatile storage device (e.g., a storage device such as read-only memory (ROM)) or memory (e.g., random access memory (RAM), memory chip, device, or integrated circuit) or an optically or magnetically readable medium (e.g., a compact disk (CD), a digital versatile disc (DVD), a magnetic disk, or a magnetic tape, etc.), regardless of whether it is erasable or rewritable. It will be understood that the storage device and the storage media are various embodiments of a non-transient machine-readable storage device suitable for storing computer programs or computer programs that include instructions that implement various embodiments of the present disclosure at execution. Accordingly, various embodiments provide a program including code for implementing the device or method claimed in one of the claims of this specification and a non-transient machine-readable storage device for storing such program.
[0371] Although the present disclosure has been illustrated and described with reference to various embodiments, those skilled in the art will understand that various changes in form and detail are possible without departing from the spirit and scope of the present disclosure as defined by the appended claims and equivalents.
Claims
1. In an electronic device (101), A housing (200) forming the exterior of the above electronic device (101); A display (250) disposed on the first surface of the housing (200); A camera (220) positioned in the direction facing the above display (250); Microphone (210); At least one processor (120) including a processing circuit; and It includes a memory (130) for storing instructions, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, A voice agent executed by the electronic device (101) operates in a conversational mode that interacts with voice input received through the microphone (210), and After receiving a first voice input, check whether at least one condition for outputting a first response related to the first voice input is satisfied, and Based on the satisfaction of at least one condition for outputting the first response, the first response associated with the first voice input is output, and Based on the fact that at least one condition for outputting the first response is not satisfied, instead of outputting the first response, causing to receive a second voice input following the first voice input, and The at least one condition for outputting the first response includes an operation of confirming that the user's gaze of the electronic device (101), obtained using the camera (220), moves to the screen of the display (250). Electronic device (101).
2. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Based on the satisfaction of the second condition, instead of outputting the first response, cause to receive the second voice input following the first voice input, and The second condition above includes an operation to confirm that the user's gaze, obtained using the camera (220), moves outside the screen of the display (250). Electronic device (101).
3. In Paragraph 1 or 2, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, After receiving the second voice input without outputting the first response, check whether at least one condition for outputting a second response related to the first voice input and the second voice input is satisfied, and Based on the satisfaction of at least one condition for outputting the second response, the second response related to the first voice input and the second voice input is output, and Based on the fact that the at least one condition for outputting the second response is not satisfied, instead of outputting the second response, causing to receive a third voice input following the second voice input, Electronic device (101).
4. In any one of paragraphs 1 to 3, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Based on confirming that a subsequent voice input is not received for a specified period from the time when the first voice input is received, checking whether the at least one condition for outputting the first response is satisfied, and Based on the satisfaction of at least one condition for outputting the first response, causing the first response associated with the first voice input to be output, Electronic device (101).
5. In any one of paragraphs 1 to 4, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Checking the continuous gaze of the user toward the screen of the display (250) from before receiving the first voice input until after receiving the first voice input, and While confirming the continuous gaze of the user toward the screen of the display (250), based on confirming that a subsequent voice input is not received for a specified period from the time the first voice input is received, causing the first response associated with the first voice input to be output. Electronic device (101).
6. In any one of paragraphs 1 through 5, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, At least based on the activation of the above conversation mode, the camera (220) is activated, and Based on the above conversation mode being disabled, causing the camera (220) to be disabled, Electronic device (101).
7. In any one of paragraphs 1 through 6, The at least one condition for outputting the first response includes an operation of verifying the designation assigned to the voice agent in the first voice input. Electronic device (101).
8. In any one of paragraphs 1 through 7, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, A point corresponding to the user's eye in an image obtained using the camera (220) is set as a tracking point, and By tracking the location of the tracking point corresponding to the user's eye, the user's gaze is confirmed, and Based on confirming the gaze of another person other than the aforementioned user, causing the aforementioned other gaze to be ignored, Electronic device (101).
9. In any one of paragraphs 1 through 8, The at least one condition for outputting the first response includes an operation of confirming that the first distance between the electronic device (101) and the user is greater than or equal to a reference distance. Electronic device (101).
10. In any one of paragraphs 1 through 9, It includes at least one sensor (231) configured to identify the posture and / or movement of the electronic device (101), and When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Using the sensing value of at least one sensor (231), the positioning state of the electronic device (101) is determined, wherein the positioning state includes a floor state in which the back surface of the electronic device (101) is placed substantially parallel to the floor surface, a standing state in which the back surface is placed at a certain angle with the floor surface, and a handheld state in which the electronic device (101) is held by a user. The at least one condition for outputting the first response includes an operation of checking whether, in the floor state, the first voice input satisfies a first condition regarding whether it includes switching information that causes the operation mode of the conversation mode to be switched, and The at least one condition for outputting the first response includes an operation to check whether the first condition or the second condition regarding the user's gaze direction is satisfied in the standing state, and The at least one condition for outputting the first response includes an operation to check whether, in the gripping state, the first condition, the second condition, or the third condition regarding the first distance between the electronic device (101) and the user is satisfied. Electronic device (101).
11. In any one of paragraphs 1 through 10, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Based on confirming the standing state or the gripping state, by activating the camera (220) of the electronic device (101), image data necessary for determining the second condition is obtained, and Causing the camera (220) to be disabled based on a transition to the floor state, failure of tracking the gaze direction, or termination of the conversation mode, Electronic device (101).
12. In any one of paragraphs 1 to 11, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Based on confirming the above gripping state, an operation to measure the first distance between the electronic device (101) and the user is performed, and Causing the measurement of the first distance to stop based on the transition to the floor state or the standing state, or the termination of the conversation mode, Electronic device (101).
13. In any one of paragraphs 1 through 12, When the above instructions are executed individually or collectively by the at least one processor (120), the electronic device (101) causes, Based on confirming a second event that causes entry into a group mode in which multiple users participate in a conversation, entering the group mode, and In order to assign titles corresponding to each of the plurality of users in the above group mode, check the input for the titles corresponding to the plurality of users, and In the group mode above, based on verifying at least one of the titles corresponding to the plurality of users in the first voice input, instead of outputting the first response, causing the receiving of the second voice input following the first voice input, Electronic device (101).
14. In the method of operating the electronic device (101), An operation in which a voice agent executed by the electronic device (101) interacts with voice input received through a microphone (210) and operates in a conversational mode, and After receiving a first voice input, an operation to check whether at least one condition for outputting a first response related to the first voice input is satisfied, and An operation of outputting the first response associated with the first voice input based on the satisfaction of at least one condition for outputting the first response, and Based on the fact that the at least one condition for outputting the first response is not satisfied, the operation of receiving a second voice input following the first voice input instead of outputting the first response is included. The at least one condition for outputting the first response includes an operation of confirming that the user's gaze of the electronic device (101), obtained using a camera (220), moves to the screen of the display (250). method.
15. In a non-transitory computer-readable recording medium for storing instructions, the instructions cause the electronic device (101) to perform at least one operation when executed individually or collectively by at least one processor (120) of the electronic device (101), and The above at least one operation is, An operation in which a voice agent executed by the electronic device (101) interacts with voice input received through a microphone (210) and operates in a conversational mode, and After receiving a first voice input, an operation to check whether at least one condition for outputting a first response related to the first voice input is satisfied, and An operation of outputting the first response associated with the first voice input based on the satisfaction of at least one condition for outputting the first response, and Based on the fact that the at least one condition for outputting the first response is not satisfied, the operation of receiving a second voice input following the first voice input instead of outputting the first response is included. The at least one condition for outputting the first response includes an operation of confirming that the user's gaze of the electronic device (101), obtained using a camera (220), moves to the screen of the display (250). Recording media.
Citation Information
Patent Citations
Device, method and program for dialog with user
JP2008217444A
Hands free device with directional interface
KR1020170013264A
Pharmaceutical composition for enhancing hyperthermia for improving gut microbiota
KR1020240046153A
Face recognition system for easy registration
KR102480910B1
Floating structure for solar panel
KR102655035B1