Electronic device and method for extracting function on basis of conversation content using same
By integrating cameras and microphones in wearable devices to recognize speakers and extract functions from conversation content, the device addresses the challenge of limited interaction in augmented reality, enhancing user convenience and efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-21
Smart Images

Figure KR2025017274_21052026_PF_FP_ABST
Abstract
Description
Electronic device and a method for extracting functions based on conversation content using the same
[0001] An embodiment of the present disclosure relates to an electronic device and a method for extracting functions based on conversation content using the same.
[0002] With the recent development of digital technology, various types of electronic devices (user devices) capable of communication and personal information processing (e.g., mobile communication terminals, PDAs (Personal Digital Assistants), electronic notebooks, smartphones, tablets, wearable electronic devices, and / or PCs (Personal Computers)) are being released. For example, electronic devices are gradually evolving into wearable electronic devices that can be worn on parts of the body to improve portability or user accessibility.
[0003] Wearable electronic devices may include head-mounted display (HMD) devices, such as glasses, that can be worn on the head. For example, wearable electronic devices may include glasses-shaped augmented reality (AR) glasses and / or smart glasses that display various content on transparent glass (e.g., lenses). As another example, wearable electronic devices may include a video see-through (VST) device that is an HMD device, captures the real environment using a camera, and displays the captured video by overlaying it onto a virtual image. Wearable electronic devices, HMD devices, and / or VST devices may use a camera to provide virtual reality services and / or augmented reality services (e.g., augmented reality worlds, augmented reality functions) to the user. For example, while the HMD device is worn on the user's head, the HMD device may implement virtual reality and / or augmented reality in response to the execution of an augmented reality-related application on a communication-connected electronic device, and may provide virtual reality services and / or augmented reality services to the user.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] An electronic device (e.g., a head-mounted display (HMD) device, a wearable electronic device) can provide an augmented reality service (e.g., a virtual reality service) based on augmented reality technology (e.g., VR (virtual reality), AR (augmented reality), MR (mixed reality), XR (extended reality)) to a user. The electronic device can implement a virtual environment interface within an augmented reality domain and can display said virtual environment interface through a display. The electronic device can provide a virtual environment interface implemented as a visual image to a user wearing said electronic device. The virtual environment interface may be implemented by mixing real objects in a real environment and virtual objects in a virtual environment.
[0006] A user can view a virtual environment interface implemented by an electronic device (e.g., an HMD device) while wearing the electronic device on a part of their body (e.g., their head). For example, the user can view a virtual environment interface in which virtual objects and real objects are mixed together or, at least partially, overlapped. The electronic device (e.g., an HMD device) may include a camera for capturing the external environment (e.g., the user's field of view), a display for showing the virtual environment interface, and a microphone for receiving external audio signals.
[0007] According to one embodiment, an electronic device (e.g., an HMD device) can acquire an audio signal (e.g., speech content) through a microphone while being worn on a part of a user's body (e.g., head), and can recognize a speaker (e.g., a conversation partner) using a camera. Based on the acquired speech content, the electronic device can extract a function related to the speech content and provide the user with a virtual environment interface containing virtual content for performing said function.
[0008] The technical tasks intended to be accomplished in this document are not limited to those mentioned above, and other technical tasks not mentioned will be clearly understood by those skilled in the art to which this document belongs from the description below.
[0009] According to one embodiment, an electronic device wearable on the head by a user may include at least one camera, a housing, a display supported on the housing and outputting visual information, a plurality of microphones for acquiring audio signals, a communication module including a communication circuit, a processor including a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by the processor, the electronic device may, in response to the acquisition of the audio signal through the plurality of microphones while the electronic device is worn on a part of the user's body, acquire first image information including a first speaker based on the at least one camera, extract a function based on the first image information and the audio signal including a conversation with the first speaker, display virtual content corresponding to the function through the display in response to the extraction of the function, and perform the function in response to input for the virtual content.
[0010] According to one embodiment, the electronic device may include a display for outputting visual information, a plurality of microphones for acquiring audio signals, a communication module including a communication circuit, a processor including a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by the processor, the electronic device may, in response to the acquisition of the audio signal through the plurality of microphones, identify a first speaker corresponding to the audio signal, extract a function based on the audio signal containing the content of a conversation with the first speaker, display virtual content corresponding to the function through the display in response to the extraction of the function, and perform the function in response to input for the virtual content.
[0011] According to one embodiment, a method for recognizing a speaker in an electronic device wearable on the head by a user may include: acquiring first image information containing a first speaker based on at least one camera in response to acquiring an audio signal through a plurality of microphones while the electronic device is worn on a part of the user's body; extracting a function based on the audio signal containing the first image information and a conversation content with the first speaker; displaying virtual content corresponding to the function through a display in response to the extraction of the function; and performing the function in response to an input to the virtual content.
[0012] According to one embodiment, a non-transient computer-readable storage medium (or computer program product) storing one or more programs for performing a method of recognizing a speaker in an electronic device wearable on the head by a user may be described. According to one embodiment, the one or more programs may include, when executed by a processor of an electronic device, an operation of acquiring first image information including a first speaker based on at least one camera in response to acquiring an audio signal through a plurality of microphones while the electronic device is worn on a part of the user's body; an operation of extracting a function based on the audio signal including the first image information and the content of a conversation with the first speaker; an operation of displaying virtual content corresponding to the function through a display in response to the extraction of the function; and an operation of performing the function in response to an input to the virtual content.
[0013] According to one embodiment, the electronic device may include an HMD device worn on a part of a user's body (e.g., head) and may provide an augmented reality service (e.g., a virtual reality service) to the user. The electronic device may implement a virtual environment interface within an augmented reality area and may recognize external objects based on the virtual environment interface.
[0014] According to one embodiment, an electronic device can recognize an external speaker and acquire the content of the conversation spoken by the speaker. The electronic device can convert the acquired conversation content into text and can visually display the conversation content. The electronic device can implement the conversation content as visual content, such as in the form of a speech bubble, and can display the visual content according to the position of the speaker. A user wearing the electronic device can intuitively identify the speaker and the conversation content. When conversing with a speaker (e.g., a conversation partner), the user may experience improved convenience in communication.
[0015] According to one embodiment, an electronic device can identify a function corresponding to context information (e.g., meeting appointment, schedule determination) based on image information including a speaker and audio information corresponding to conversation content, and can provide the identified function to a user. While conversing with the speaker, the user can identify a function related to the conversation content and efficiently utilize the function. The electronic device can provide an efficient user experience to the user.
[0016] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.
[0017] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. The aforementioned features and advantages will be clearly understood based on the attached drawings and the description thereof.
[0018] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment of the present disclosure.
[0019] FIG. 2a is a perspective view schematically showing the configuration of a first type of XR device supporting an XR (extended reality) service according to one embodiment of the present disclosure.
[0020] FIG. 2b is a perspective view schematically showing the front of a second type of XR device according to one embodiment of the present disclosure.
[0021] FIG. 2c is a perspective view schematically showing the rear side of a second type of XR device according to one embodiment of the present disclosure.
[0022] FIG. 3 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0023] FIG. 4 is an exemplary diagram showing processing operations between components according to one embodiment of the present disclosure.
[0024] FIG. 5a is a flowchart illustrating a method for extracting functions based on conversation content according to one embodiment of the present disclosure.
[0025] FIG. 5b is a flowchart illustrating a method for displaying virtual content based on a previous conversation record with a speaker according to one embodiment of the present disclosure.
[0026] FIG. 6a is an example diagram for identifying a speaker according to one embodiment of the present disclosure.
[0027] FIG. 6b is an example illustration reflecting a visual emphasis effect on an identified speaker according to one embodiment of the present disclosure.
[0028] FIG. 6c is an example diagram showing the content of speech by a speaker as visual content according to one embodiment of the present disclosure.
[0029] FIG. 6d is an example diagram showing a summary of a speech content summarized according to one embodiment of the present disclosure.
[0030] FIG. 6e is an example diagram showing the entire utterance content in response to user input for a summary according to one embodiment of the present disclosure.
[0031] FIG. 7a is a first example diagram illustrating the execution of a function corresponding to situation information according to one embodiment of the present disclosure.
[0032] FIG. 7b is a second example diagram illustrating the execution of a function corresponding to situation information according to one embodiment of the present disclosure.
[0033] FIG. 7c is a third exemplary diagram illustrating the execution of a function corresponding to situation information according to one embodiment of the present disclosure.
[0034] FIG. 8a is a first example diagram illustrating that when a function corresponding to situation information according to one embodiment of the present disclosure is identified, the speaker is matched with the identified function, and the speaker and the function are stored in speech-related information.
[0035] FIG. 8b is a second example diagram in which, when a function corresponding to situation information according to one embodiment of the present disclosure is identified, the speaker is matched with the identified function, and the speaker and the function are stored in speech-related information.
[0036] FIG. 8c is a first example diagram for identifying a speaker and a function matched to the speaker based on speech-related information according to one embodiment of the present disclosure.
[0037] FIG. 8d is a second example diagram for identifying a speaker and a function matched to the speaker based on speech-related information according to one embodiment of the present disclosure.
[0038] FIG. 9 is an example diagram showing a first utterance content corresponding to a first speaker and a second utterance content corresponding to a second speaker in a situation where a conversation is having with a plurality of speakers according to one embodiment of the present disclosure.
[0039] FIG. 10 is an example diagram showing how a portable electronic device according to one embodiment of the present disclosure provides a function to a user based on context information based on conversation content.
[0040] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0041] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0042] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0043] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0044] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0045] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0046] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0047] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0048] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0049] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0050] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0051] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0052] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0053] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0054] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0055] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0056] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0057] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0058] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0059] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0060] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0061] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0062] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0063] FIG. 2a is a perspective view schematically showing the configuration of a first type of XR device (201) supporting an XR (extended reality) service according to one embodiment of the present disclosure. FIG. 2b is a perspective view schematically showing the front of a second type of XR device (2000) according to one embodiment of the present disclosure. FIG. 2c is a perspective view schematically showing the rear of a second type of XR device (2000) according to one embodiment of the present disclosure.
[0064] In FIG. 2a, for example, a first type of XR device (201) (e.g., the electronic device (101) of FIG. 1) is illustrated as being implemented as glasses-type AR glasses, but this is merely an example and can be replaced with an electronic device that supports XR (extended reality) services or immersive media services. For example, the XR device (201) of FIG. 2a may include the XR device (2000) of FIG. 2b and FIG. 2c.
[0065] Referring to FIG. 2a, an electronic device (101) according to one embodiment may be an XR device that provides images related to an XR (extended reality) service to a user (e.g., the XR device (201) of FIG. 2a, or the XR device (2000) of FIG. 2b and FIG. 2c). XR or XR service may be defined as a service that collectively refers to virtual reality (VR), augmented reality (AR), or mixed reality (MR). The XR device (201) may be a head-mounted device (HMD), a head-mounted display (HMD), or AR glasses worn on the user's body or head, but is not limited thereto. FIG. 2a illustrates, by way of example, that the XR device (201) is implemented as glasses-type AR glasses, but this is merely an example. The XR device (2000) illustrated in FIG. 2b and FIG. 2c may also be an exemplary electronic device of a different form. The XR device (201) of FIG. 2a and the XR device (2000) of FIG. 2b and 2c can provide images related to XR services to the user.
[0066] The XR device (201) may include an OST (optical see-through) type configured to allow external light to reach the user's eyes through glass when worn, or a VST (video see-through) type configured to allow light emitted from a display to reach the user's eyes when worn, while blocking external light so that external light does not reach the user's eyes.
[0067] The XR device (201) can provide a user with an image related to an extended reality (XR) service. For example, the XR device (201) can provide an XR service image (or immersive media) in which at least one virtual content (e.g., an image, video, or object) is displayed in a field of view (FoV) region determined to be the user's field of view. The field of view determined to be the user's field of view is an area determined to be perceptible to the user through the XR device (201), and may be an area that includes all or at least part of the display of the XR device (201).
[0068] The XR device (201) may further include at least some of the configurations and / or functions of FIG. 1, and in the case of configurations that overlap with FIG. 1, they may be substantially the same configuration.
[0069] For example, the XR device (201) may include a display unit (210), a transparent member (220), an optical member (230), a camera unit (240), a lighting unit (250), a microphone unit (260), a speaker unit (270), a battery unit (280), and a PCB unit (290), but this is merely an example and is not limited thereto. The XR device (201) may include a housing containing the aforementioned components, and each component may be at least partially supported or fixed to the housing. For example, the XR device (201) may include a glass-shaped housing.
[0070] The display unit (210) may include a first display (211) placed on the right eye and a second display (212) placed on the left eye. The parts / components constituting the first display (211) and the second display (212) may be the same. For example, the display unit (210) may be referred to as including a first display (211), a second display (212), a transparent member (220), a screen display unit (2210) (e.g., combiner optics), a display driving unit (213), and a lens (e.g., a projection lens) (not shown), but this is merely an example.
[0071] The first display (211) and the second display (212) may include a liquid crystal on silicon (LCoS), a light emitting diode (LED) on silicon (LEDoS), an organic light emitting diode (OLED), a micro LED, and / or a digital mirror device (DMD). The first display (211) and the second display (212) may be composed of a display area composed of pixels for displaying an image and light receiving pixels (e.g., photo sensor pixels) placed between the pixels that receive light reflected from the eye, convert it into electrical energy, and output it.
[0072] A transparent member (220) is positioned on the front of the display (211, 212) and can protect the display (211, 212). The transparent member (220) may be formed from a glass plate, a plastic plate, or a polymer, and may be made transparent or translucent. The transparent member (220) can control the transmission of light incident on the display (211, 212). The transparent member (220) or the screen display (2210) may include a lens including a waveguide (e.g., a waveguide) and / or a reflective lens.
[0073] Light emitted from the first display (211) and the second display (212) can be transmitted to the user's eye by passing through a lens (not shown) and a waveguide (e.g., a waveguide) and reflecting off the input optical member (222) and the grating area formed on the screen display unit (2210). The waveguide may be made of glass, plastic, and / or polymer and may include a nano-pattern formed on one surface, for example, a polygonal or curved grating structure. The waveguide may include at least one diffractive element, such as a DOE (diffractive optical element), HOE (holographic optical element), or a reflective element (e.g., a reflective mirror). In one embodiment, the waveguide may guide the light emitted from the displays (211, 212) to the user's eye using at least one diffractive element or reflective element included in the waveguide. For example, in the case of an AR device (201) that provides augmented reality, the user can perceive the actual space (or actual environment) behind the display (211, 212) by passing through the display (211, 212).
[0074] The camera unit (240) may include a first camera (241), a second camera (242), and a third camera (243). The first camera (241) may be referred to as HR (high resolution) or PV (photo video) and may include a high-resolution camera. The first camera (211) may include a color camera equipped with functions for obtaining high-quality images, such as AF (auto focus) and shake correction (OIS (optical image stabilizer)). The first camera (241) may include a GS (global shutter) camera or an RS (rolling shutter) camera. The second camera (242) and the third camera (243) can perform at least one of 3DoF (3 degrees of freedom), 6DoF head tracking, hand detection and tracking, pose estimation and prediction, gesture recognition, slam function through depth capture, spatial recognition and / or eye tracking.
[0075] For example, the second camera (242) and the third camera (243) may be implemented as at least one of a stereo camera for head tracking and spatial recognition, a gesture camera for detecting the user's movement (or hand), an eye tracking camera for tracking the direction of gaze by tracking the movement of the user's left eye and / or right eye, a distance measuring camera (e.g., TOF (time of flight camera) or depth camera) for measuring the distance to an object located in front of the XR device (201), a SLAM camera (simultaneous localization and mapping camera) for recognizing information related to the surrounding space (e.g., location and / or direction), and / or an RGB camera for detecting color-related information of an object and distance information to an object, or may be implemented as a camera in which some of these are integrated.
[0076] The lighting unit (250) may be placed in the housing (e.g., frame) of the XR device (201). The lighting unit (250) may be used as a means to supplement ambient brightness when taking a picture with a camera. For example, the lighting unit (250) may be used when subject detection is not easy due to at least one of the following environments: a dark environment, an environment where multiple light sources are mixed, or an environment where reflected light occurs. As another example, when the detection of the pupil is not easy, the lighting unit (250) may irradiate infrared wavelengths to facilitate the detection of the pupil.
[0077] The microphone unit (260) includes a plurality of microphones and can process external acoustic signals into electrical voice data. The processed voice data can be utilized in various ways depending on the function (or application running) being performed on the XR device (201). The speaker unit (270) includes a plurality of speakers and can output audio data under the control of the processor (120).
[0078] The battery unit (280) can supply power to drive components (e.g., the processor (120), memory (130), sensor module (176) and communication module (190) shown in FIG. 1) placed in the display unit (210), camera unit (240), lighting unit (250), microphone unit (260), speaker unit (270) and PCB unit (290).
[0079] The sensor module (176) may include various sensors for detecting measurements related to the movement of the XR device (201) (e.g., velocity, acceleration, angular velocity, angular acceleration, and / or geographic location). For example, the sensor module (176) may include, but is not limited to, a proximity sensor, an ambient light sensor, a geomagnetic sensor, a gesture sensor, an accelerometer, and / or a gyroscope. The proximity sensor may detect objects adjacent to the XR device (201). The ambient light sensor may measure the degree of brightness around the XR device (201). The gyroscope may detect the state (or posture, orientation) and position of the XR device (201). The gyroscope may detect the movement of the XR device (201) or a user wearing the XR device (201). For example, the XR device (201) can use an illuminance sensor to check the brightness level around the XR device (201) and change the brightness-related setting information of the display unit (210) based on the brightness level.
[0080] The memory (130) can store instructions that can be executed by the processor (120) or the electronic device (101). The processor (120) can execute the instructions stored in the memory (130) to implement a software module and can control the components included in the XR device (201). The operation of the processor (120) described below can be performed when the instructions stored in the memory (130) are executed.
[0081] The processor (120) acquires XR service image information related to a space corresponding to the field of view of a user wearing the XR device (201) and can recognize an area (hereinafter, FOV) determined to be the user's field of view (FOV). The processor (120) can control the display unit (210) so that at least a portion of the XR service image is displayed through the recognized FOV. For example, a user wearing the XR device (201) can perceive an external environment (e.g., background image) seen through the display unit (210) and a virtual object (e.g., virtual content) output through the display unit (210) as being at least partially mixed. The user can distinguish and recognize the actual external environment image and the virtual object.
[0082] The processor (120) can measure values related to the movement of the XR device (201) (e.g., velocity, acceleration, angular velocity, angular acceleration, and geographic location), and can obtain movement information of the XR device (201) (e.g., rotation angle) using the measured values or a combination thereof. The processor (120) can analyze the movement information of the XR device (201) and XR service video information in real time and control the processing of actions required to provide XR services, e.g., head tracking and / or eye tracking actions.
[0083] According to one embodiment, the XR device (201) can provide an XR service video (or immersive media) in which at least one virtual content (e.g., image, video, or object) is displayed based on a field of view (FoV) region determined to be the user's field of vision. The XR device (201) can implement a virtual environment interface and output it through a display unit (210), and can process various tasks within the virtual environment interface. The XR device (201) can recognize information related to tasks occurring within the virtual environment interface by time or by space (location), and can store it in the form of data blocks. According to one embodiment, when loading a specific task that was processed in the past, the XR device (201) can display the specific task by time or by space.
[0084] According to one embodiment, the XR device (201) can implement a time interface that appears separated by time and a spatial interface that appears separated by space in relation to a specific task, and can implement a virtual environment interface in which the time interface and the spatial interface are at least partially integrated. According to one embodiment, when the XR device (201) provides an augmented reality service to a user, the convenience of the user using the augmented reality service can be improved.
[0085] The electronic device illustrated in FIGS. 2b and 2c may be a second type of extended reality (XR) device (2000). In one embodiment, the XR device (2000) may be a device that provides augmented reality (e.g., services related to augmented reality) to a user. In one embodiment, the XR device (2000) may be a virtual reality (VR) device that provides virtual reality to a user. For example, the XR device (2000) may include a video see-through (VST) device. The XR device (2000) of FIGS. 2b and 2c may be included in the XR device (201) of FIG. 2a.
[0086] According to one embodiment, as illustrated in FIGS. 2b and 2c, the XR device (2000) may include a distance sensor (2500), a face recognition camera (2310, 2320), a printed circuit board (not shown), a first display module (3010) and / or a second display module (3020). In some embodiments, the XR device (2000) may be implemented by including at least some of the components included in the electronic device (101) of FIG. 1, or by additionally including other components. The location or form of the components included in the XR device (2000) may be varied and is not limited to the examples illustrated in FIGS. 2b and 2c.
[0087] According to one embodiment, the XR device (2000) may include a housing (2100). In one embodiment, the housing (2100) may include a first surface (2110) (e.g., front) exposed to an external environment and a second surface (2120) (e.g., rear) that is not exposed to an external environment and adheres to the user's skin when worn. For example, when the XR device (2000) is worn on the user's face, the first surface (2110) of the XR device (2000) may be exposed to an external environment, and the second surface (2120) of the XR device (2000) may be in contact with the user's face at least partially. In one embodiment, the XR device (2000) may adhere to the user's face through various components. For example, the XR device (2000) may use a band formed of an elastic material attached to the housing (2100) to make the second surface (2120) of the housing (2100) adhere to the area around the user's eyes. In another embodiment, the XR device (2000) may be worn on the user's face via eyeglass temples, helmets, or straps. Additionally, the XR device (2000) may be partially worn on the user's face through various configurations.
[0088] In one embodiment, the housing (2100) of the XR device (2000) may be formed to have a shape or structure that is easy to wear on a user's face. For example, the housing (2100) may have a second surface (2120) formed in a streamlined shape so as to cover part of the user's eyes and nose. In one embodiment, a nose-shaped recess (2130) may be formed on the second surface (2120) of the housing (2100) so as to be supported by the user's nose.
[0089] In one embodiment, the housing (2100) of the XR device (2000) may be formed of a lightweight material (e.g., plastic) that allows the user to feel a comfortable fit. Meanwhile, the housing (2100) may be formed of a non-metallic and / or metallic material having a certain level of rigidity against external impact. The metallic material may include aluminum, stainless steel (STS, SUS), iron, magnesium, or alloys such as titanium. The non-metallic material may include synthetic resin, ceramic, or engineering plastic.
[0090] According to one embodiment, as illustrated in FIG. 2b and FIG. 2c, at least one first camera module (2410, 2420) (e.g., camera module (190) of FIG. 1) may be positioned in the front direction of the XR device (2000) (e.g., the -Y direction with respect to FIG. 2b, the direction of the user's gaze). For example, the XR device (2000) may include a camera module (2410) corresponding to the user's left eye and a camera module (2420) corresponding to the user's right eye. The XR device (2000) may capture the external environment in the front direction of the XR device (2000) (e.g., the -Y direction with respect to FIG. 2b) using at least one first camera module (2410, 2420).
[0091] According to one embodiment, at least one first camera module (2410, 2420) illustrated in FIG. 2b may include one or more lenses, an image sensor, and / or an image signal processor. In one embodiment, the location or number of at least one first camera module (2410, 2420) may vary and is not limited to the illustrated example. In one embodiment, at least one first camera module (2410, 2420) may measure depth of field (DOF). The XR device (2000) can perform various functions such as head tracking, hand detection or tracking, gesture recognition, or spatial recognition using a depth of field (e.g., 3DOF (degrees of freedom) or 6DOF) obtained through at least one first camera module (2410, 2420). The at least one first camera module (2410, 2420) may include, for example, a GS (global shutter) camera or an RS (rolling shutter) camera, and the location or number thereof may vary and is not limited to the illustrated examples. According to some embodiments, the at least one first camera module (2410, 2420) can recognize the surrounding space of the XR device (2000). The at least one first camera module (2410, 2420) can detect a user's gesture within a certain distance (e.g., a certain space) of the XR device (2000). At least one first camera module (2410, 2420) may include a GS (global shutter) camera that can reduce the RS (rolling shutter) phenomenon in order to detect and track fast hand movements and / or fine movements of the user's fingers.
[0092] According to one embodiment, as illustrated in FIG. 2b, at least one second camera module (2210, 2220, 2230 and / or 2240) may be disposed on a first surface (2110) of a housing (2100). In one embodiment, the second camera module (2210, 2220, 2230, 2240) may acquire image data for an external image. The external image data obtained through at least one second camera module (2210, 2220, 2230, 2240) may be transmitted to the user through a display module (3000) disposed on the user's left eye and right eye, respectively. The location or number of at least one second camera module (2210, 2220, 2230 and / or 2240) may vary and is not limited to the example illustrated in FIG. 2b.
[0093] According to one embodiment, as illustrated in FIG. 2b, at least one distance sensor (2500) may be disposed on a first surface (2110) of the housing (2100). For example, the at least one distance sensor (2500) may measure the distance to at least one object disposed around the XR device (2000). The at least one distance sensor (2500) may include an infrared sensor, an ultrasonic sensor, and / or a light detection and ranging (LiDAR) sensor. The at least one distance sensor (2500) may be implemented based on an infrared sensor, an ultrasonic sensor, and / or a LiDAR sensor. The location or number of distance sensors (2500) may vary and is not limited to the example illustrated in FIG. 2b.
[0094] According to one embodiment, as illustrated in FIG. 2c, the XR device (2000) may include an eye tracking camera module (4100). The eye tracking camera module (4100) can detect and track the user's pupils. The eye tracking camera module (4100) can track the user's gaze or the direction of the user's head using at least one method among, for example, an EOG sensor (electro-oculography or electrooculogram), a coil system, a dual Purkinje system, bright pupil systems, or dark pupil systems. Additionally, the eye tracking camera module (4100) may include a GS (global shutter) camera to track the user's rapid pupil movements.
[0095] In one embodiment, the eye-tracking camera module (4100) may include at least one camera unit (4110) (e.g., a micro camera or an IR LED) for tracking the wearer's gaze, for example, by being positioned on a second surface (2120) of the housing (2100). In one embodiment, the eye-tracking camera module (4100) may include a first eye-tracking camera module (4100-1) for tracking the user's left eye and a second eye-tracking camera module (4100-2) for tracking the user's right eye, which are positioned on the second surface (2120) of the housing (2100). The XR device (2000) can determine the direction in which the user is looking based on the movement of the pupils tracked using a plurality of eye-tracking camera modules (4100-1, 4100-2). The XR device (2000) can determine the direction of the user's head using a plurality of eye-tracking camera modules (4100-1, 4100-2).
[0096] The term "eye-tracking camera module (4100)" used in various embodiments of the present disclosure may be used interchangeably with terms such as the third camera module.
[0097] According to one embodiment, the XR device (2000) can detect the eye corresponding to the dominant eye and / or auxiliary eye among the user's left eye and / or right eye by using a first eye-tracking camera module (4100-1) and / or a second eye-tracking camera module (4100-2). For example, the XR device (2000) can detect the eye corresponding to the dominant eye and / or auxiliary eye based on the user's gaze toward an external object or a virtual object and / or the direction of the user's head.
[0098] According to one embodiment, as illustrated in FIG. 2c, at least one face recognition camera (2310, 2320) may be disposed on the second surface (2120) of the XR device (2000). For example, a plurality of face recognition cameras (2310, 2320) may recognize the user's face when the XR device (2000) is worn on the user's face. In one embodiment, the face recognition cameras (2310, 2320) may detect the user's facial expression. In one embodiment, the XR device (2000) may determine whether the XR device (2000) is worn on the user's face using a plurality of face recognition cameras (2310, 2320). In one embodiment, the camera unit (4110) of the eye-tracking camera module (4100) may be a face recognition camera.
[0099] According to one embodiment, as illustrated in FIG. 2b, the XR device (2000) may include at least one light-emitting member (4200). For example, the light-emitting member (4200) may provide state information of the XR device (2000) in the form of light. As another example, the light-emitting member (4200) may provide a light source that is coupled with the operation of the eye-tracking camera module (4100). The light-emitting member (4200) may include, for example, an LED, an IR LED, or a xenon lamp. In one embodiment, the light-emitting member (4200) may emit light to increase the accuracy of the first eye-tracking camera module (4100-1), the second eye-tracking camera module (4100-2), the first face recognition camera (2310), and / or the second face recognition camera (2320).
[0100] According to one embodiment, with reference to FIG. 2b, the light-emitting member (4200) may include a first light-emitting member (4200-1) disposed in a first display module (3010) and a second light-emitting member (4200-2) disposed in a second display module (3020). The first light-emitting member (4200-1) emits light to the user's left eye, thereby increasing accuracy when the first eye-tracking camera module (4100-1) photographs the user's left eye. The second light-emitting member (4200-2) emits light to the user's right eye, thereby increasing accuracy when the second eye-tracking camera module (4100-2) photographs the user's right eye.
[0101] According to one embodiment, the XR device (2000) may include a plurality of display modules (3000) disposed on a second surface (2120) located in the rear direction of the XR device (2000) (e.g., +Y direction with respect to FIG. 2c, opposite direction of the user's gaze direction). For example, a first display module (3010) corresponding to the user's left eye and a second display module (3020) corresponding to the user's right eye may be disposed on the second surface (2120) of the XR device (2000). For example, when the XR device (2000) is worn on the user's face, the first display module (3010) may be disposed corresponding to the user's left eye, and the second display module (3020) may be disposed corresponding to the user's right eye.
[0102] FIG. 3 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0103] The electronic device (101) of FIG. 3 may be at least partially similar to the electronic device (101) of FIG. 1, the XR device (201) of FIG. 2a (e.g., a wearable electronic device) and / or, the XR device (2000) of FIG. 2b and FIG. 2c, or may further include other embodiments of the electronic device (101) and the XR device (201, 2000). According to one embodiment, the electronic device (101) may include at least partially similar components to the XR device (201) of FIG. 2a. For example, the electronic device (101) may include one of a head-mounted device (HMD), a head-mounted display (HMD), or AR glasses worn on the user's body or head as shown in FIG. 2a, and may include an XR device (201) that provides images related to an extended reality (XR) service to the user.
[0104] Referring to FIG. 3, the electronic device (101) may include a processor (120) (e.g., processor (120) of FIG. 1), a memory (130) (e.g., memory (130) of FIG. 1), a display (160) (e.g., display module (160) of FIG. 1, display unit (210) of FIG. 2a), a camera (180) (e.g., camera module (180) of FIG. 1, camera unit (240) of FIG. 2a), a microphone (310) (e.g., input module (150) of FIG. 1), a speaker (320) (e.g., sound output module (155) of FIG. 1), and / or a communication circuit (190) (e.g., communication module (190) of FIG. 1). In the memory (130) of the electronic device (101), an augmented reality-related application (331) for implementing an augmented reality space (e.g., a virtual environment interface), information related to an artificial intelligence model (332) for analyzing situational information, and speech-related information (333) including a speaker and speech content matched to the speaker may be stored. According to one embodiment, various information related to the augmented reality space and virtual content (e.g., virtual objects, speech bubbles, visual effects) may be stored in the memory (130).
[0105] According to one embodiment, a processor (120) of an electronic device (101) can execute a program stored in memory (130) (e.g., program (140) of FIG. 1, augmented reality related application (331)) to control at least one other component (e.g., hardware and / or software component) and perform various data processing or operations. According to one embodiment, the processor (120) may include at least one processor including a processing circuit. The number of processors (120) may be one or more. For example, the processor (120) may include the structure of a multi-core processor such as a dual core, a quad core, and / or a hexa core.
[0106] According to one embodiment, the memory (130) may store instructions (e.g., commands) that are executed individually or collectively by the processor (120) (e.g., at least one processor). The processor (120) may control the operation of components included in the electronic device (101) by executing the commands stored in the memory (130). For example, the processor (120) may include a plurality of processors and may control the operation of a plurality of operations to be divided individually among the plurality of processors or to be performed collectively. According to one embodiment, the processor (120) may be operatively, functionally, and / or electrically connected to the memory (130), display (160), camera (180), microphone (310), speaker (320), and / or communication module (190).
[0107] According to one embodiment, the processor (120) can execute an augmented reality-related application (331) stored in memory (130) and can implement an augmented reality space (e.g., a virtual environment interface) based on the augmented reality-related application (331). The processor (120) can display the augmented reality space (e.g., a virtual environment interface) through a display (160). According to one embodiment, the processor (120) can automatically execute the augmented reality-related application (331) in response to the situation where the electronic device (101) is worn on a user's head and can provide an augmented reality service to the user. The user can view the virtual environment interface displayed through the display (160) of the electronic device (101).
[0108] According to one embodiment, an electronic device (101) may be operatively connected to a dedicated controller for providing augmented reality services to a user within an augmented reality space. For example, the dedicated controller may include input means for controlling an augmented reality-related application (331) at least partially. For example, the dedicated controller may include an input device implemented in hardware (e.g., a controller, a remote control, a stylus pen), or may include an input module and input means implemented in software (e.g., visually) within the augmented reality space. As another example, the electronic device (101) may recognize a user's hand (e.g., a finger) as the dedicated controller and may control an augmented reality-related application (331) based on input and movements using the user's hand. According to one embodiment, a processor (120) may detect user input (e.g., prompt input, audio input, touch input, and / or gesture input) and movements based on the dedicated controller and may perform functions and / or actions related to augmented reality based on said user input. For example, the processor (120) can detect an XR event signal for capturing a virtual environment image based on an augmented reality service. The processor (120) can capture the virtual environment image in response to the detection of the XR event signal.
[0109] According to one embodiment, the processor (120) can execute various functions and programs related to augmented reality within an augmented reality space. For example, the processor (120) can execute a program for processing specific tasks (e.g., schedule registration, schedule management, execution of specific applications) within the augmented reality space and can display virtual objects based on said program. For example, when a program for schedule registration (e.g., a calendar application, a schedule-related application) is executed, the processor (120) can display a virtual page for schedule registration and provide a virtual environment in which a user can efficiently register a schedule.
[0110] According to one embodiment, the processor (120) can analyze situational information based on artificial intelligence model-related information (332) stored in memory (130) and can extract functions corresponding to said situational information. For example, the situational information may be identified based on at least one of image information captured using a camera (180) and audio information obtained through a microphone (310). For example, the image information may include location information, status information, and / or feature information about the other party (e.g., speaker) when the user is talking to the electronic device (101). The audio information may include the content of speech (e.g., conversation content) of the other party (e.g., speaker) and / or phoneme information of the other party. According to one embodiment, the artificial intelligence model-related information (332) may include an artificial intelligence model (e.g., an AI model) for analyzing situational information identified based on image information and audio information.
[0111] According to one embodiment, the processor (120) can identify the content of a conversation between a user and a counterpart (e.g., speaker), contextual information regarding the user, and / or contextual information regarding the counterpart based on artificial intelligence model-related information (332), and can extract a function corresponding to said contextual information. For example, in a situation where the user and the counterpart are having a conversation related to a meeting schedule (e.g., time, date, day of the week), the processor (120) can extract a “function related to schedule registration.” In another example, in a situation where the user and the counterpart are having a conversation related to a meeting place, the processor (120) can extract a “function related to finding a place.” The processor (120) can generate virtual content (e.g., a virtual icon, an execution icon for executing a specific application) to execute the extracted function, and can display said virtual content through a display (160). The processor (120) can execute a specific application related to the extracted function in response to user input regarding the virtual content. According to one embodiment, the processor (120) can extract a function corresponding to situational information based on artificial intelligence model-related information (332), and can provide virtual content for the execution of the function to the user so that the extracted function can be executed intuitively. While conversing with another person, the user can check virtual content (e.g., a function) related to the conversation content and efficiently execute a function corresponding to the virtual content.
[0112] According to one embodiment, the processor (120) can identify the content of a conversation between a user and a counterpart (e.g., a speaker), situational information regarding the user, and / or situational information regarding the counterpart based on artificial intelligence model-related information (332), and can also generate a summary corresponding to the conversation content based on the artificial intelligence model-related information (332). In a situation where the conversation content becomes long (e.g., a situation where the time for inputting the conversation content exceeds a set threshold), the processor (120) can generate a summary based on the conversation content and display the generated summary through a display (160).
[0113] According to one embodiment, the processor (120) can identify a counterpart (e.g., speaker, conversation partner), conversation content (e.g., utterance content) matching the counterpart (e.g., utterance content), and / or a function extracted based on the conversation content, based on utterance-related information (333) stored in memory (130). For example, the utterance-related information (333) may include conversation content with the counterpart at a past point in time. The utterance-related information (333) may store the counterpart and the conversation content with the counterpart in response to a situation where a specific function is extracted based on the conversation content. The utterance-related information (333) may store the counterpart and the conversation content mapped to the counterpart together.
[0114] According to one embodiment, the camera (180) may be positioned so that the direction of the lens faces the front of the electronic device (101) (e.g., the direction of the user's gaze) and may capture the actual environment. The camera (180) may function as an input module for acquiring the actual environment corresponding to the user's field of vision. For example, the processor (120) may use the camera (180) to capture the external environment and the conversation partner corresponding to the front of the electronic device (101), and may acquire a captured image (e.g., image information) containing the external environment and the conversation partner. The processor (120) may acquire territory information (e.g., location feature information, characteristic information, background information) and feature information about the partner based on the captured image. The processor (120) may store the captured image (e.g., image information) in memory (130).
[0115] According to one embodiment, the camera (180) may include a pupil tracking camera for tracking the user's pupil movement (e.g., gaze direction). For example, the processor (120) may use the pupil tracking camera to determine the user's gaze direction and identify a conversation partner (e.g., speaker) based on the gaze direction. The pupil tracking camera may function as an input module for obtaining information related to the user's gaze direction.
[0116] According to one embodiment, the display (160) can visually display programs running and activated functions on the electronic device (101). For example, when an augmented reality-related application (331) is executed, the processor (120) can implement an augmented reality space (e.g., a virtual environment interface) and display the augmented reality space through the display module (160). According to one embodiment, the electronic device (101) can visually display an augmented reality space implemented based on an augmented reality service to a user through the display module (160). In one embodiment, a user wearing the electronic device (101) on their head may feel as if they are living in a real-world space.
[0117] According to one embodiment, the communication circuit (190) can perform a communication connection with an external electronic device (e.g., the electronic device (102, 104) of FIG. 1). For example, the electronic device (101) (e.g., HMD device) and the external electronic device (102, 104) (e.g., portable electronic device) can be connected via the communication module (190) according to various communication methods (e.g., wired communication channel, wireless communication channel).
[0118] According to one embodiment, the microphone (310) may function as an input module for acquiring an external audio signal (e.g., voice signal). For example, the processor (120) may use the microphone (310) to acquire speech content (e.g., conversation content) by a conversation partner (e.g., speaker). According to one embodiment, the microphone (310) may include a directional microphone positioned to face the speaker, taking into account the position of the speaker (e.g., conversation partner).
[0119] According to one embodiment, the electronic device (101) can independently implement an augmented reality space and provide augmented reality services to a user. According to one embodiment, the electronic device (101) may be at least partially controlled by the external electronic device (102) while being operatively connected to the external electronic device (102) (e.g., mobile electronic device, smartphone) through a communication circuit (190). For example, when an augmented reality space is implemented in the external electronic device (102), the electronic device (101) may acquire the augmented reality space implemented from the external electronic device (102) and may display the acquired augmented reality space through a display (160).
[0120] According to one embodiment, a processor (120) of an electronic device (101) (e.g., an HMD device) can detect a situation in which the electronic device (101) is worn on a user's head and can execute an augmented reality-related application (311) stored in memory (130). Based on the executed augmented reality-related application (311), the processor (120) can implement a virtual environment interface and display the implemented virtual environment interface through a display (160).
[0121] According to one embodiment, an electronic device (101) may acquire image information captured using a camera (180) in response to the acquisition of an audio signal (e.g., speech content) through a microphone (310) within a virtual environment interface. For example, a user wearing the electronic device (101) may move their gaze toward the direction in which the audio signal is generated, and the shooting direction of the camera (180) corresponding to the user's gaze direction may change. The user may be in a situation where they are looking at a conversation partner who generated the audio signal (e.g., speech content). The image information captured using the camera (180) may include the surrounding environment and the conversation partner. The electronic device (101) may analyze the image information captured using the camera (180) and the audio signal (e.g., audio information) acquired using the microphone (310). The image information and the audio information may be included in the context information. For example, the electronic device (101) can analyze the situation information based on the artificial intelligence model-related information (332) stored in the memory (130) and can extract a function corresponding to the situation information. By analyzing the situation information, the electronic device (101) can extract a function required in the current situation and can display virtual content for performing the function through the display (160). The electronic device (101) can perform a function corresponding to the virtual content in response to user input regarding the virtual content.
[0122] According to one embodiment, an electronic device (101) can identify a function corresponding to context information (e.g., meeting appointment, schedule determination) based on image information including a speaker (e.g., conversation partner) and audio information corresponding to conversation content, and can provide the identified function to a user. While conversing with the speaker, the user can visually identify a function related to the conversation content (e.g., a function required in the current situation during the conversation) and can intuitively execute the function on the electronic device (101). The electronic device (101) can provide an efficient user experience to the user.
[0123] According to one embodiment, an electronic device (101) wearable on the head by a user may include at least one camera (180), a housing, a display (160) supported on the housing and outputting visual information, a plurality of microphones (310), a communication circuit (190) including a communication circuit, a processor (120) including a processing circuit, and a memory (130) for storing instructions. When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) can acquire an audio signal through the plurality of microphones (310) while the electronic device (101) is worn on a part of the user's body, acquire first image information based on the at least one camera (180) in response to the acquisition of the audio signal, extract a function corresponding to situation information based on the audio signal corresponding to the first image information and the utterance content of the first speaker included in the first image information, display virtual content for executing the function through the display (160) in response to the extraction of the function, and execute the function in response to input for the virtual content.
[0124] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can identify a past conversation record for the first speaker and a function execution record mapped to the past conversation record based on speech-related information (333) stored in the memory (130), and can extract the function associated with the first speaker based on the identified past conversation record and the identified function execution record.
[0125] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can update the utterance-related information (333) corresponding to the first utterer when performing the function.
[0126] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) may display virtual content corresponding to the first speaker based on the identified past conversation record, and in response to input for the virtual content, display the identified past conversation record.
[0127] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) may generate the first summary for the first speaker based on the past conversation record for the first speaker in response to the conversation time with the first speaker exceeding a set threshold time, and display the original text corresponding to the first summary in response to an input to the first summary.
[0128] According to one embodiment, the electronic device (101) may further include a pupil tracking camera for detecting the user's pupil movement. When the instructions are executed individually or collectively by the processor (120), the electronic device (101) may recognize the first speaker that the user is gazing at based on the pupil movement identified using the pupil tracking camera, and display a visual highlighting effect along the outside of the recognized first speaker.
[0129] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can adjust the reception direction of the plurality of microphones (310) toward the recognized first speaker, and receive the audio information containing the content of the conversation with the first speaker based on the plurality of microphones (310) with the adjusted reception direction.
[0130] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can identify the utterance order and the flow of conversation content for the plurality of speakers based on the first image information, and based on the utterance order and the flow of conversation content, display virtual content individually for each of the plurality of speakers and display a visual highlighting effect on the virtual content corresponding to the speaker who is speaking.
[0131] According to one embodiment, the function may include at least one of registering a schedule, transferring a file, searching for a location, and executing a set application.
[0132] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can identify the location of the first speaker based on the first image information, determine the display location of the virtual content based on the location of the first speaker, and display the virtual content according to the determined display location.
[0133] According to one embodiment, an electronic device (e.g., a portable electronic device) may include a display for outputting visual information, a plurality of microphones for acquiring audio signals, a communication module including a communication circuit, a processor including a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by the processor, the electronic device may, in response to the acquisition of the audio signal through the plurality of microphones, identify a first speaker corresponding to the audio signal, extract a function based on the audio signal containing the content of a conversation with the first speaker, display virtual content corresponding to the function through the display in response to the extraction of the function, and perform the function in response to input for the virtual content.
[0134] FIG. 4 is an exemplary diagram showing processing operations between components according to one embodiment of the present disclosure.
[0135] According to one embodiment, each operation illustrated in FIG. 4 may be understood to be performed by a processor (e.g., processor (120) of FIG. 1 and FIG. 3, processing circuit, at least one processor) of an electronic device (e.g., electronic device (101) of FIG. 1 and FIG. 3). The electronic device (101) of FIG. 4 may include at least some similarities to the XR device (201) of FIG. 2a, the XR device (2000) of FIG. 2b and FIG. 2c, and the electronic device (101) of FIG. 3, or may include other embodiments of the electronic device (101). For example, the electronic device (101) may include an XR device (201), an electronic device in the form of glasses, and / or a wearable electronic device, as illustrated in FIG. 2a. The electronic device (101) may include a wearable electronic device (e.g., a head-mounted display (HMD) device) that operates while being worn at least partially on a part of the user's body (e.g., the head).
[0136] Referring to FIG. 4, the electronic device (101) can signal process (420) data input through the input module (410) and output the signal-processed data through the output module (430). For example, the first data input through the input module (410) may include image data and audio data. The second data output through the output module (430) may include visually displayed content.
[0137] According to one embodiment, the input module (410) may include a microphone (411) for acquiring an audio signal (e.g., audio information, audio data, and / or voice data), a camera (412) for acquiring a captured image (e.g., image information), and / or a pupil tracking camera (413) for tracking the user's pupil movement (e.g., gaze direction). For example, the camera (412) may be positioned to photograph an external environment corresponding to the user's field of vision. The pupil tracking camera (413) may be positioned to photograph the user's pupil.
[0138] According to one embodiment, the processor (120) can acquire the content of a speaker's speech (e.g., speech voice data (421), audio information) using a microphone (411). The processor (120) can acquire a captured image (e.g., image information) corresponding to the external environment (e.g., real environment) that the user is looking at using a camera (412). For example, the captured image may include a conversation partner (e.g., speaker) with whom the user is conversing. The processor (120) can acquire speaker data (422) included in the captured image. The processor (120) can monitor the user's pupil movement (e.g., gaze direction) using a pupil tracking camera (413) and identify the object the user is looking at (e.g., conversation partner). The processor (120) can acquire speaker data (422) corresponding to the user's gaze direction. Referring to FIG. 4, in the signal processing (420) process, the processor (120) can apply speech voice data (421) and speaker data (422) to an artificial intelligence model (423). For example, the artificial intelligence model (423) may include an AI algorithm contained in artificial intelligence model related information (332) of a memory (130) (e.g., memory (130) of FIG. 3).
[0139] According to one embodiment, the processor (120) can apply spoken voice data (421) (e.g., audio information) and speaker data (422) (e.g., image information) to an artificial intelligence model (423) (e.g., artificial intelligence model related information (332)), and can analyze the spoken voice data (421) and the speaker data (422). For example, the spoken voice data (421) and the speaker data (422) may be included in context information. The processor (120) can analyze context information based on the situation in which a user and a conversation partner are conversing, and can extract a function corresponding to the context information. For example, the processor (120) can extract a function corresponding to the context information based on at least one word included in the spoken voice data (421). The extracted function may include at least one of a function to register a schedule, a function to transfer a file, and / or a function to execute a specific application.
[0140] For example, if it is confirmed that a meeting schedule is being decided during a conversation between a user and a conversation partner, the processor (120) can extract a function for registering the schedule. If it is confirmed that a file (e.g., a photo) is being transmitted during a conversation between a user and a conversation partner, the processor (120) can extract a function for transmitting the file. If it is confirmed that a meeting place is being decided during a conversation between a user and a conversation partner, the processor (120) can extract a function for searching for a place. According to one embodiment, the processor (120) can determine a specific application for performing the extracted functions. For example, a calendar application can perform the function of registering a schedule, a gallery application can perform the function of transmitting a photo, and a map application can perform the function of searching for a place. The processor (120) can determine a specific application corresponding to each function and can generate virtual content (e.g., a virtual icon, a virtual item) for executing the specific application.
[0141] According to one embodiment, the output module (430) can generate virtual content (431) for executing a specific application and can output the virtual content through a display (432) (e.g., the display (160) of FIG. 3).
[0142] According to one embodiment, the processor (120) can execute a specific application corresponding to the virtual content in response to user input (e.g., interaction input) for the virtual content.
[0143] FIG. 5a is a flowchart illustrating a method for extracting functions based on conversation content according to one embodiment of the present disclosure. FIG. 5b is a flowchart illustrating a method for displaying virtual content based on a previous conversation record with a speaker according to one embodiment of the present disclosure.
[0144] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0145] According to one embodiment, operations 501 to 511 of FIG. 5a and operations 521 to 527 of FIG. 5b may be understood to be performed by a processor (e.g., processor (120) of FIG. 1 and FIG. 3, processing circuit, at least one processor) of an electronic device (e.g., electronic device (101) of FIG. 1 and FIG. 3). The electronic device of FIG. 5a and FIG. 5b may include at least some similarities to the XR device (201) of FIG. 2a, the XR device (2000) of FIG. 2b and FIG. 2c, and the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). For example, the electronic device (101) may include an XR device (201), an electronic device in the form of glasses, and / or a wearable electronic device, as shown in FIG. 2a. The electronic device (101) may include a wearable electronic device (e.g., a head-mounted display (HMD) device) that operates while being worn at least partially on a part of the user's body (e.g., the head).
[0146] Referring to FIG. 5a, in operation 501, the processor (120) of the electronic device (101) can execute a virtual environment program (e.g., an augmented reality-related application (331) of FIG. 3) stored in memory (e.g., memory (130) of FIG. 3) in response to the situation where the electronic device (101) is worn on a part of the user's body (e.g., head). For example, the electronic device (101) can detect the situation where it is worn on the user's head and can automatically execute the virtual environment program. As another example, the electronic device (101) may execute the virtual environment program by user input.
[0147] In operation 503, the processor (120) can acquire spoken voice data (e.g., audio information, audio signal) through a microphone (e.g., microphone (310) of FIG. 3). For example, when a virtual environment program is executed, the microphone (310) may be at least partially activated, and the processor (120) can acquire external spoken voice data using the activated microphone (310). The spoken voice data may include audio signals generated by a conversation partner (e.g., speaker) who is conversing with the user. The spoken voice data may include conversation content spoken by the conversation partner.
[0148] According to one embodiment, the processor (120) can convert acquired speech voice data (e.g., audio signal) into text in the form of characters and can include and display the text in a speech bubble. For example, the processor (120) can identify the location (e.g., coordinate information) of the conversation partner (e.g., speaker), create a virtual speech bubble for the conversation partner, include the text inside the virtual speech bubble, and display the virtual speech bubble in a form adjacent to the face location of the conversation partner. For example, the processor (120) can create a virtual speech bubble containing the speech content of the conversation partner and display the virtual speech bubble for the conversation partner.
[0149] In operation 505, the processor (120) can acquire first image information (e.g., a captured image) containing the speaker based on a camera (e.g., the camera (180) of FIG. 3). For example, in operation 505, a user wearing the electronic device (101) may move their gaze toward the direction in which spoken voice data is generated (e.g., the location of the conversation partner) and may be in a state of looking at the speaker.
[0150] In operation 507, the processor (120) can identify a function corresponding to context information based on spoken voice data (e.g., audio information) and a first image information. For example, context information may include audio information corresponding to spoken voice data and image information corresponding to the first image information. According to one embodiment, the processor (120) can analyze context information based on an artificial intelligence model (e.g., information related to the artificial intelligence model in FIG. 3 (332)) and extract a function corresponding to the context information. The processor (120) can extract a function that matches the user's intention in accordance with the situation in which the user and the speaker (e.g., conversation partner) are conversing. For example, if a situation is confirmed in which the user and the speaker decide on a meeting schedule, the processor (120) can determine that the user wants to register the schedule and can extract a function for registering the schedule (e.g., an application related to schedule registration, a calendar application). As another example, if it is confirmed that the user sends a photo to the speaker, the processor (120) can extract a function for sending a file (e.g., an application related to file transfer, a gallery application). As another example, if it is confirmed that the user and the speaker decide on a meeting place, the processor (120) can extract a function for searching for a place (e.g., an application related to place search, a route finding application).
[0151] In operation 509, the processor (120) may display virtual content for executing a identified function. For example, the virtual content may include an execution icon (e.g., execution item, execution object) for executing a specific application (e.g., an application determined based on context information). The processor (120) may display the virtual content through the display (160) while the user and the conversation partner are conversing.
[0152] In operation 511, the processor (120) can execute a function corresponding to the virtual content in response to user input (e.g., interaction input) for the virtual content.
[0153] According to one embodiment, an electronic device (101) can extract a function corresponding to situation information in a situation where a user and a conversation partner are conversing, and can generate and display virtual content corresponding to said function to directly execute said extracted function. For example, a user can receive virtual content to perform a function intended by the user even while conversing with a conversation partner. The electronic device (101) can directly execute a function corresponding to said virtual content in response to user input regarding said virtual content.
[0154] According to one embodiment, when an electronic device (101) confirms a record of a previous conversation (e.g., past conversation content) with a conversation partner (e.g., speaker), it can generate and display virtual content based on the past conversation content. Operations 521 to 527 of FIG. 5b may be operations that describe in detail operations 507 and 509 of FIG. 5a.
[0155] Referring to FIG. 5b, in operation 521, the processor (120) of the electronic device (101) can identify a speaker (e.g., conversation partner) based on spoken voice data and first image information. For example, the processor (120) can identify a speaker based on the characteristics of the speaker included in the spoken voice data and the first image information, and can determine whether the identified speaker is a speaker with whom a conversation has previously taken place. The processor (120) can identify past conversation content (e.g., past conversation record, past conversation history) corresponding to the speaker based on speech-related information (333) stored in memory (130).
[0156] In operation 523, the processor (120) can determine whether it has conversed with the speaker in the past. For example, the processor (120) can check whether past conversation content about the speaker is stored in memory (130), and if the past conversation content is stored in memory (130), it can determine that it has conversed with the speaker in the past.
[0157] When past conversation content with the speaker is confirmed, in operation 525, the processor (120) may display virtual content related to the past conversation content. For example, the virtual content may provide the user with the entire content of the past conversation content in response to user input. Referring to FIGS. 5a and 5b, the processor (120) may display a first virtual content to perform a specific function while displaying a second virtual content to confirm the past conversation content with the speaker. The processor (120) may display the first virtual content and the second virtual content together, or may display the first virtual content and the second virtual content sequentially in response to a set input.
[0158] In operation 527, the processor (120) can map current speech voice data (e.g., current conversation content) to the speaker and store it in speech-related information (333) in memory (130). For example, if the speaker is a first speaker with whom the conversation has previously taken place (e.g., a state where past conversation content is stored in memory (130)), the processor (120) can update the speech-related information (333) corresponding to the first speaker. For example, if the speaker is a second speaker with whom the conversation has not previously taken place (e.g., a state where there is no past conversation content), the processor (120) can map the second speaker and speech voice data for the second speaker and store it in speech-related information (333).
[0159] According to one embodiment, the electronic device (101) can check past conversation content with the conversation partner in a situation where the user and the conversation partner are conversing, and if the past conversation content exists, it can display virtual content for outputting the past conversation content. For example, while conversing with the conversation partner, the user can check past conversation content with the conversation partner, and communication with the conversation partner can be smooth. The electronic device (101) can directly display past conversation content with the conversation partner in response to user input regarding the virtual content. It can immediately provide the user with past conversation content regarding the conversation partner.
[0160] FIG. 6a is an exemplary diagram for identifying a speaker according to one embodiment of the present disclosure. FIG. 6b is an exemplary diagram reflecting a visual emphasis effect on the identified speaker according to one embodiment of the present disclosure. FIG. 6c is an exemplary diagram for displaying the content of a speech by a speaker as visual content according to one embodiment of the present disclosure. FIG. 6d is an exemplary diagram for displaying a summary of the content of the speech according to one embodiment of the present disclosure. FIG. 6e is an exemplary diagram for displaying the entire content of the speech in response to user input regarding the summary according to one embodiment of the present disclosure.
[0161] The electronic device of FIGS. 6a through 6e may be at least partially similar to the XR device (201) of FIG. 2a, the XR device (2000) of FIG. 2b and 2c, and the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). The electronic device (101) may include a wearable electronic device (e.g., a head-mounted display (HMD) device) that operates while being worn at least partially on a part of a user's body (e.g., the head). The operation in FIGS. 6a through 6e may be understood as being performed by a processor of the electronic device (101) (e.g., the processor (120) of FIG. 3, a processing circuit, at least one processor).
[0162] FIG. 6a is an exemplary diagram for identifying a speaker according to one embodiment of the present disclosure. FIG. 6a is an image (e.g., screen) that a user wearing an electronic device (101) on their head can see through the display of the electronic device (101) (e.g., the display (160) of FIG. 3). The electronic device (101) can use a camera (e.g., the camera (180) of FIG. 3) to acquire a captured image corresponding to the user's field of vision and can display the captured image through the display (160). For example, the user can identify a first counterpart (601) and a second counterpart (602) in an actual environment.
[0163] FIG. 6b is an example illustration in which a visual emphasis effect is reflected for an identified speaker according to one embodiment of the present disclosure. FIG. 6b is a situation in which the first counterpart (601) speaks among the first counterpart (601) and the second counterpart (602). A processor (120) of an electronic device (101) can identify a speaker (e.g., the first counterpart (601)) in response to a situation in which an audio signal (e.g., speech voice data) is acquired. For example, the processor (120) can detect the movement of the mouth of the first counterpart (601) and confirm that the first counterpart (601) is in a state of speaking. To indicate that the first counterpart (601) is the speaker, the processor (120) can display a visual emphasis effect (e.g., a highlight effect) based on the first counterpart (601). For example, the processor (120) may display a highlight line (603) along the outline (e.g., border) of the first counterpart (601). The user may recognize that the first counterpart (601) with the highlight line (603) displayed is the speaker.
[0164] Referring to FIG. 6b, the processor (120) can recognize that the first counterpart (601) is the speaker in response to the acquisition of an audio signal and can adjust the receiving direction of the microphone (e.g., microphone (310) in FIG. 3, a plurality of microphones) so that the receiving direction of the microphone is directed toward the speaker (e.g., the first counterpart (601)). For example, the microphone (310) may include a plurality of directional microphones and the receiving direction may be adjusted to be directed toward a specific direction. The electronic device (101) can at least partially adjust the receiving direction of the microphone (310) so that the receiving direction of the microphone (310) is directed toward the first counterpart (601) in response to confirming that the speaker is the first counterpart (601). The electronic device (101) can receive the speech of the first counterpart (601) based on the microphone (310) with the adjusted receiving direction.
[0165] According to one embodiment, the electronic device (101) can track the direction of a user's gaze and recognize a first counterpart (601) corresponding to the direction of the gaze. For example, the processor (120) can display a visual emphasis effect in relation to the recognized first counterpart (601). In another example, if the processor (120) recognizes that the first counterpart corresponding to the direction of the gaze is speaking, it can display a visual emphasis effect for the first counterpart (601) who is the speaker.
[0166] FIG. 6c is an example illustration of displaying speech content by a speaker as visual content (e.g., a speech bubble) according to one embodiment of the present disclosure. FIG. 6c is a situation in which speech voice data for a first counterpart (601) is converted into text, and a speech bubble (611) containing the converted text is displayed. For example, a processor (120) can acquire the speech voice data through a microphone (e.g., the microphone (310) of FIG. 3) and can convert the speech voice data into text.
[0167] According to one embodiment, the electronic device (101) may display the converted text in the form of a speech bubble in response to a situation in which the first counterpart (601) is recognized as a speaker (e.g., a situation in which the user's gaze direction is toward the first counterpart (601)). For example, the processor (120) may obtain the location of the first counterpart (601) (e.g., the first counterpart (601) recognized as a speaker) in the actual environment (604) as first coordinate information. Based on the first coordinate information (e.g., location information of the first counterpart (601)), the processor (120) may determine the location where the speech bubble (611) is to be displayed (e.g., second coordinate information). For example, the processor (120) may determine the location where the speech bubble (611) is to be displayed based on the locations of various objects included in the captured image and the location of the first counterpart (601) (e.g., first coordinate information). The speech bubble (611) can be displayed adjacent to the first counterpart (601). For example, if the spoken voice data is “Did you see the news on TV right now?, I think a wanted criminal was on the news, but doesn’t that face look familiar?”, the sentences can be converted into text, and the converted text can be included in the speech bubble (611).
[0168] According to one embodiment, the electronic device (101) may display the utterance as virtual content (e.g., speech bubble (611)) or output an audio signal corresponding to the utterance through a speaker. For example, if the electronic device (101) is in a state of communication connection with an external electronic device (e.g., audio output device, wireless earphones), the electronic device (101) may at least partially control the external electronic device so that an audio signal corresponding to the utterance is output based on the external electronic device.
[0169] FIG. 6d is an exemplary illustration showing a summary of speech content according to one embodiment of the present disclosure. FIG. 6d shows a situation in which a speech bubble (611) containing a summary (612) of speech voice data (e.g., text of speech bubble (611) in FIG. 6c) of a first counterpart (601) and additional text (613) for additional speech voice data is displayed. When the length (e.g., capacity) of the speech voice data exceeds a set threshold, the processor (120) may generate a summary (612) by summarizing the speech voice data. For example, the summary (612) may include text summarized based on the speech voice data and an option button (612-1) for checking the full content of the speech voice data. Additional text (613) may be generated based on additional speech voice data that is not included in the summary (612). For example, the summary (612) of FIG. 6d can be generated based on the text contained in the speech bubble (611) of FIG. 6c. The summary (612) can be displayed as the text “While talking about the news, I saw a wanted criminal.” Additional text (613) can include the text “Don’t you think you’ve seen it too? Someone here looks like that person.”
[0170] FIG. 6e is an example diagram showing the entire speech content in response to user input for a summary according to one embodiment of the present disclosure. FIG. 6e is a situation in which the entire content (614) of the spoken voice data is displayed in a speech bubble in response to input of the option button (612-1) in FIG. 6d. For example, the entire content (614) may include text contained in the speech bubble (611) of FIG. 6c and additional text (613) of FIG. 6d.
[0171] FIG. 7a is a first exemplary diagram illustrating the execution of a function corresponding to situation information according to one embodiment of the present disclosure. FIG. 7b is a second exemplary diagram illustrating the execution of a function corresponding to situation information according to one embodiment of the present disclosure. FIG. 7c is a third exemplary diagram illustrating the execution of a function corresponding to situation information according to one embodiment of the present disclosure.
[0172] The electronic device of FIGS. 7a through 7c may be at least partially similar to the XR device (201) of FIG. 2a, the XR device (2000) of FIG. 2b and 2c, and the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). The electronic device (101) may include a wearable electronic device (e.g., a head-mounted display (HMD) device) that operates while being worn at least partially on a part of a user's body (e.g., the head). The operation in FIGS. 7a through 7c may be understood as being performed by a processor of the electronic device (101) (e.g., the processor (120) of FIG. 3, a processing circuit, at least one processor).
[0173] FIG. 7a is an image (e.g., screen) that a user wearing the electronic device (101) on their head sees through the display of the electronic device (101) (e.g., the display (160) of FIG. 3). The processor (120) of the electronic device (101) can identify a speaker (e.g., a conversation partner (701)) in response to the situation in which an audio signal (e.g., spoken voice data) is acquired. The processor (120) can display a visual emphasis effect (702) (e.g., a highlight effect) to indicate the speaker (701). The processor (120) can acquire spoken voice data of the speaker (701) through a microphone (e.g., the microphone (310) of FIG. 3), convert said spoken voice data into text, and display a speech bubble (711) containing said converted text. For example, the speech bubble (711) may contain the utterance “Are you free next Tuesday at 7 o’clock? Let’s have a meal at a restaurant we frequent.”
[0174] FIG. 7b is an example of extracting functions based on context information. For example, context information may include image information being displayed through the display (160) and speech content (e.g., audio information) contained in the speech bubble (711) of FIG. 7a. The processor (120) can analyze context information based on an artificial intelligence model (e.g., information related to the artificial intelligence model (331) of FIG. 3) and extract functions related to the topic of the speech content. Referring to FIG. 7b, the processor (120) can recognize that the situation is one in which the user and the speaker (701) decide on an “appointment” (e.g., schedule). The processor (120) can extract functions related to the “appointment” (e.g., an application related to schedule registration) and display virtual content (712) for executing said functions. For example, the virtual content (712) may include a guidance message for executing said functions. For example, the guidance message may include the content, “Schedule-related content has been recognized in the utterance. Would you like to register it on the calendar?” For example, virtual content (712) may include an icon, item, object, and / or function shortcut object for performing a specific function (e.g., registering a schedule). When user input for virtual content (712) is received, the processor (120) may perform a specific function corresponding to said virtual content (712).
[0175] FIG. 7c is an example of executing a function (e.g., schedule registration) in response to user input (e.g., interaction input) for virtual content (712). The processor (120) can execute a calendar application (e.g., an application related to schedule registration) corresponding to the function (e.g., schedule registration) in response to user input for the virtual content (712) of FIG. 7b, and can display an execution screen (713) of the calendar application. The processor (120) can directly execute the calendar application in response to a single user input and display an execution screen (713) for schedule registration.
[0176] According to one embodiment, the electronic device (101) can extract a function (e.g., an application related to schedule registration) related to the content of the conversation (e.g., content regarding deciding on an appointment) while the user is conversing with the speaker (701), and can display virtual content (712) for performing the extracted function. The electronic device (101) can display an execution screen (713) for schedule registration in response to input regarding the virtual content (712). The electronic device (101) can directly provide the user with a function that meets the user's needs. The electronic device (101) can provide the user with an efficient user experience.
[0177] According to one embodiment, if the conversation content is related to an appointment, the electronic device (101) may determine a schedule-related application for registering the appointment. According to one embodiment, if the conversation content describes an image or is related to a work of art, the electronic device (101) may determine at least one application among a gallery application for viewing an image, a drawing-related application for drawing an image, and / or a search-related application for searching for a work of art. According to one embodiment, if the conversation content is related to work, the electronic device (101) may determine a schedule-related application for managing a work schedule. According to one embodiment, if the conversation content includes a contact, the electronic device (101) may determine a call-related application for contacting the contact and / or a message-related application for sending a message to the contact.
[0178] According to one embodiment, the electronic device (101) can analyze the conversation content of the speaker (701) based on an artificial intelligence model (e.g., information related to the artificial intelligence model in FIG. 3 (331)), and can extract functions that meet the needs of the user of the electronic device (101) based on the analysis results. The electronic device (101) can directly provide the functions and operations desired by the user, and the user experience utilizing the electronic device (101) can be improved.
[0179] FIG. 8a is a first exemplary diagram in which, when a function corresponding to situation information according to one embodiment of the present disclosure is identified, the speaker is matched with the identified function, and the speaker and the function are stored in speech-related information. FIG. 8b is a second exemplary diagram in which, when a function corresponding to situation information according to one embodiment of the present disclosure is identified, the speaker is matched with the identified function, and the speaker and the function are stored in speech-related information.
[0180] FIG. 8c is a first exemplary diagram illustrating the identification of a speaker and a function matched to the speaker based on speech-related information according to one embodiment of the present disclosure. FIG. 8d is a second exemplary diagram illustrating the identification of a speaker and a function matched to the speaker based on speech-related information according to one embodiment of the present disclosure.
[0181] The electronic device of FIGS. 8a through 8d may be at least partially similar to the XR device (201) of FIG. 2a, the XR device (2000) of FIG. 2b and 2c, and the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). The electronic device (101) may include a wearable electronic device (e.g., a head-mounted display (HMD) device) that operates while being worn at least partially on a part of a user's body (e.g., the head). The operation in FIGS. 8a through 8d may be understood as being performed by a processor of the electronic device (101) (e.g., the processor (120) of FIG. 3, a processing circuit, at least one processor).
[0182] FIG. 8a is an image (e.g., screen) that a user wearing the electronic device (101) on their head sees through the display of the electronic device (101) (e.g., the display (160) in FIG. 3). The processor (120) of the electronic device (101) can identify a speaker (e.g., a conversation partner (801)) in response to acquiring an audio signal (e.g., spoken voice data). The processor (120) can display a visual emphasis effect (802) (e.g., a highlight effect) to indicate the speaker (801). The processor (120) can acquire spoken voice data of the speaker (801) through a microphone (e.g., the microphone (310) in FIG. 3), convert said spoken voice data into text, and display a speech bubble (811) containing said converted text. For example, the speech bubble (811) is “Seho, go to the nearby mart and buy some eggs and green onions.” It may include speech content. For example, the conversation partner (801) may mean the user’s “mom,” and the user wearing the electronic device (101) may mean “Seho.”
[0183] FIG. 8b is an example of extracting functions based on contextual information. For example, contextual information may include image information (803) being displayed through the display (160) and speech content (e.g., audio information) contained in the speech bubble (811) of FIG. 8a. The processor (120) can analyze contextual information based on an artificial intelligence model (e.g., information related to the artificial intelligence model (331) of FIG. 3) and extract functions related to the topic of the speech content. Referring to FIG. 8b, the processor (120) can recognize that the speaker (801) (e.g., “Mom”) is giving a “running errand.” The processor (120) can extract functions related to the “running errand” (e.g., an application related to memo recording) and display virtual content (812) for executing said functions. For example, the virtual content (812) may include a guidance message for executing said functions. For example, the guidance message may include the content “A request has been detected in the utterance. Would you like to register the utterance-related information along with the speaker (e.g., “Mom”)?”. The processor (120) may respond to user input regarding virtual content (812) by mapping the speaker (801) to the extracted function (e.g., an application related to memo recording) and storing it in memory (e.g., memory (130) of FIG. 3).
[0184] According to one embodiment, the processor (120) can recognize the speaker (801) and, based on the face of the speaker (801), extract at least one feature and map the speaker (801) to the extracted at least one feature. The processor (120) can store the speaker (801) in speech-related information (e.g., speech-related information (333) of FIG. 3) based on the speaker (801) and at least one feature corresponding to the speaker (801). According to one embodiment, when a request (e.g., at least one function) corresponding to the speech content by the speaker (801) is identified, the processor (120) can map the speaker (801) to the identified request. The processor (120) can store the speaker (801) and the request corresponding to the speaker (801) in speech-related information (333).
[0185] FIG. 8c may be a situation in which an electronic device (101) recognizes a speaker (e.g., conversation partner (801), “Mom”) already stored in speech-related information (333). A processor (120) may display a visual emphasis effect (802) (e.g., a highlight effect) to indicate the speaker (801). The processor (120) may identify the speaker (801) based on a captured image (803) obtained through a camera (e.g., camera (180) in FIG. 3) and determine whether the speaker (801) matches the speaker stored in the speech-related information (333). For example, if the speaker (801) matches the speaker stored in the speech-related information (333), the processor (120) may display a notification message (821). For example, the notification message (821) may include the content “Speaker detected - Mom.” According to one embodiment, the processor (120) may, in response to the detection of a speaker included in image information (803), display information representing the speaker (e.g., “Mom”) as a notification message (821).
[0186] FIG. 8d shows that an electronic device (101) can identify a speaker (801) and a previously performed function corresponding to the speaker (801) based on speech-related information (333). When a previously performed function (e.g., a request, at least one function, or errands in FIG. 8a and FIG. 8b) mapped to the speaker (801) (e.g., “Mom”) is identified, the processor (120) can display a notification message (822) related to said function. For example, the notification message (822) may include virtual content (823) in which a request related to said function (e.g., “buy eggs and green onions at the mart”) is disclosed. For example, the virtual content (823) may include an icon, an item, and / or an object for performing a specific function (e.g., checking past schedules, checking past conversation content). According to one embodiment, when user input for virtual content (712) is received, the electronic device (101) can check the details of past conversations with the speaker (801) (e.g., “Mom”).
[0187] According to one embodiment, the electronic device (101) can identify a previously performed function (e.g., requests, past conversation content, and / or, at least one function) mapped to the speaker (801) while the user is conversing with the speaker (801), and if the previously performed function is identified, the function can be displayed as a notification message (822). The electronic device (101) can help the user remember the speaker (801) by providing the user with a past record related to the speaker (801). The electronic device (101) can provide the user with an efficient user experience.
[0188] FIG. 9 is an example diagram showing a first utterance content corresponding to a first speaker and a second utterance content corresponding to a second speaker in a situation where a conversation is having with a plurality of speakers according to one embodiment of the present disclosure.
[0189] The electronic device of FIG. 9 may include at least some similarities to the XR device (201) of FIG. 2a, the XR device (2000) of FIG. 2b and FIG. 2c, and the electronic device (101) of FIG. 3, or may include other embodiments of the electronic device (101). The electronic device (101) may include a wearable electronic device (e.g., a head-mounted display (HMD) device) that operates while being worn at least partially on a part of a user's body (e.g., the head). The operation in FIG. 9 may be understood as being performed by a processor of the electronic device (101) (e.g., the processor (120) of FIG. 3, a processing circuit, at least one processor).
[0190] FIG. 9 is an example illustration depicting a situation in which a conversation is held with multiple speakers (e.g., a first speaker (901), a second speaker (902)). FIG. 9 is an image (e.g., a screen) that a user wearing an electronic device (101) on their head checks through the display of the electronic device (101) (e.g., the display (160) of FIG. 3). The electronic device (101) can use a camera (e.g., the camera (180) of FIG. 3) to acquire a captured image corresponding to the user's field of vision and can display the captured image through the display (160). For example, the user can identify the first speaker (901) and the second speaker (902) in the actual environment.
[0191] Referring to FIG. 9, the processor (120) may display a visual first emphasis effect (903) (e.g., a highlight effect) based on the first speaker (901) in response to the situation in which the first speaker (901) speaks. For example, the first emphasis effect (903) may be displayed in a first color. The processor (120) may convert the speech content of the first speaker (901) into first text and display a first speech bubble (911) containing the first text. For example, the first speech bubble (911) may display the content “I agree with the conclusion of the meeting right now.”
[0192] Referring to FIG. 9, the processor (120) may display a visual second emphasis effect (904) (e.g., a highlight effect) based on the second speaker (902) in response to the situation in which the second speaker (902) speaks. For example, the second emphasis effect (904) may be displayed in a second color different from the first color. The processor (120) may convert the speech content of the second speaker (902) into second text and display a second speech bubble (912) containing the second text. For example, the second speech bubble (912) may display the content “According to the data I investigated, a new conclusion is reached.”
[0193] According to one embodiment, when a user wearing an electronic device (101) converses with a plurality of speakers (e.g., a first speaker (901), a second speaker (902)), the electronic device (101) may display a number in each speech bubble based on the order in which the speech content is acquired. For example, when the first speaker (901) speaks first, the processor (120) may display “Number 1” in the first-1 speech bubble corresponding to the first-1 speech content of the first speaker (901). After the first-1 speech content is acquired, when the second speaker (902) speaks next, the processor (120) may display “Number 2” in the second-1 speech bubble corresponding to the second-1 speech content of the second speaker (902). According to one embodiment, the electronic device (101) may display sequence information in a speech bubble corresponding to the utterance content according to the conversation flow and the order of acquisition of the utterance content. According to another embodiment, the electronic device (101) may display an arrow pointing to the next speech bubble based on the order of acquisition of the utterance content. For example, if a second-1 utterance content is acquired after a first-1 utterance content in the order of acquisition, the processor (120) may display a first-1 speech bubble corresponding to the first-1 utterance content, and then display the arrow such that an arrow starting from the first-1 speech bubble points to a second-1 speech bubble corresponding to the second-1 utterance content.
[0194] According to one embodiment, the electronic device (101) can display virtual content (e.g., numbers, arrows) indicating the order of speech content even when a user and multiple speakers are conversing. The user can directly understand the flow of speech content. Efficiency can be improved through the use of the electronic device (101).
[0195] FIG. 10 is an example diagram showing how a portable electronic device according to one embodiment of the present disclosure provides a function to a user based on context information based on conversation content.
[0196] The electronic device (1001) of FIG. 10 may be at least partially similar to the electronic device (101) of FIG. 1 and FIG. 3, or may include other embodiments of the electronic device (101). The electronic device (1001) may include a portable electronic device having a different form factor from the XR device (201) of FIG. 2a. Operation in FIG. 10 may be understood to be performed by a processor of the electronic device (1001) (e.g., processor (120) of FIG. 3, processing circuit, at least one processor).
[0197] Referring to FIG. 10, a user of an electronic device (1001) may be in a state of being in a conversation with a counterpart (e.g., a speaker) of another electronic device. A processor (120) of the electronic device (1001) may acquire a first speech of the user using a microphone (e.g., the microphone (310) in FIG. 3), convert the acquired first speech into a first text (1011), and display the first text (1011) through a display (e.g., the display (160) in FIG. 3). For example, the first text (1011) may include the content “Are you free next Tuesday at 7 o’clock?” The processor (120) can obtain a second utterance of a conversation partner (e.g., speaker), convert the obtained second utterance into a second text (1012), and display the second text (1012) through a display (160). For example, the second text (1012) may include the content “Let’s have a meal at a restaurant we frequent.”
[0198] According to one embodiment, the processor (120) can analyze a first text (1011) and a second text (1012) based on an artificial intelligence model (e.g., information related to the artificial intelligence model in FIG. 3 (331)) and can extract at least one function based on the analysis result. For example, the processor (120) can extract at least one function (e.g., schedule registration function, place search function, route finding function) based on a first keyword (1013) (“Tuesday 7 o’clock”) included in the first text (1011) and a second keyword (1014) (“restaurant I used to visit often”) included in the second text (1012).
[0199] For example, the processor (120) may determine a calendar application for performing a schedule registration function and may display a first execution icon (1021) corresponding to the calendar application. For example, the processor (120) may determine a map application for performing a place search function and may display a second execution icon (1022) corresponding to the map application. For example, the processor (120) may determine a navigation application for performing a route finding function and may display a third execution icon (1023) corresponding to the navigation application.
[0200] According to one embodiment, the processor (120) can execute a schedule registration function based on a first keyword (1013) and a second keyword (1014) in response to user input for a first execution icon (1021). For example, the processor (120) can display a schedule registration screen through a display (160) and can set at least one word among the first keyword (1013) and the second keyword (1014) to be entered into an input field included in the schedule registration screen. According to one embodiment, the electronic device (1001) can display a schedule registration screen tailored to the user's needs in response to a single input for the first execution icon (1021).
[0201] According to one embodiment, the processor (120) can execute a place search function based on a first keyword (1013) and a second keyword (1014) in response to user input for a second execution icon (1022). For example, the processor (120) can display a place search screen through a display (160) and can set up input fields included in the place search screen to input words related to “restaurants frequently visited” representing “places” among the first keyword (1013) and the second keyword (1014). According to one embodiment, the electronic device (1001) can display a place search screen tailored to the user's needs in response to a single input for the second execution icon (1022).
[0202] According to one embodiment, the processor (120) can execute a navigation function based on a first keyword (1013) and a second keyword (1014) in response to user input for a third execution icon (1023). For example, the processor (120) can display a navigation screen through a display (160) and can set up input fields included in the navigation screen to input words related to “restaurants frequently visited” representing “places” among the first keyword (1013) and the second keyword (1014). According to one embodiment, the electronic device (1001) can display a navigation screen tailored to the user's needs in response to a single input for a first execution icon (1021).
[0203] According to one embodiment, an electronic device (1001) can extract at least one function based on at least one keyword included in audio information in a situation where a user and a conversation partner (e.g., a speaker) are conversing, and can generate and display virtual content (e.g., icons (1021, 1022, 1023)) corresponding to said function to directly execute said extracted at least one function. For example, a user may be provided with virtual content (1021, 1022, 1023) to perform a function intended by the user even while conversing with a conversation partner. The electronic device (1001) can directly execute a function corresponding to said virtual content in response to user input regarding said virtual content (1021, 1022, 1023).
[0204] According to one embodiment, the electronic device is not limited to a specific form factor and may include various electronic devices having various form factors. For example, the electronic device may include a foldable electronic device. The foldable electronic device may be mounted or fixed in a state where a camera is facing a conversation partner (e.g., a speaker), and while maintaining such a state, the conversation partner can be identified and the conversation partner's speech content can be acquired. The foldable electronic device may convert the acquired speech content into visual content and display it. Based on the acquired speech content, the foldable electronic device may extract a function related to the speech content and provide the extracted function to the user.
[0205] A method for recognizing a speaker in an electronic device (101) that can be worn on the head by a user according to one embodiment may include: acquiring an audio signal through a plurality of microphones (310) while the electronic device is worn on a part of the user's body; acquiring first image information based on at least one camera (180) in response to the acquisition of the audio signal; extracting a function corresponding to situation information based on the audio signal corresponding to the first image information and the conversation content of the first speaker included in the first image information; displaying virtual content corresponding to the function through the display (160) in response to the extraction of the function; and performing the function in response to input for the virtual content.
[0206] A method according to one embodiment may further include an operation to check a past conversation record and a past function execution record for the first speaker based on speaker-related information stored in memory, an operation to extract the function related to the first speaker based on the checked past conversation record and the checked past function execution record, and an operation to update the speaker-related information corresponding to the first speaker when the function is performed.
[0207] A method according to one embodiment may further include an operation of displaying virtual content corresponding to the first speaker based on the identified past conversation record, and an operation of displaying the identified past conversation record in response to an input regarding the virtual content.
[0208] A method according to one embodiment may further include an operation of generating a first summary for the first speaker based on a past conversation record for the first speaker in response to the conversation time with the first speaker exceeding a set threshold time, and an operation of displaying an original text corresponding to the first summary in response to an input for the first summary.
[0209] A method according to one embodiment may further include an action of detecting a user's pupil movement using a pupil tracking camera, an action of recognizing a first speaker whom the user is gazing at based on the detected pupil movement, and an action of displaying a visual emphasis effect along the outer side of the recognized first speaker.
[0210] A method according to one embodiment may further include the operation of adjusting the reception direction of the plurality of microphones toward the recognized first speaker, and the operation of receiving the audio signal containing the content of a conversation with the first speaker based on the plurality of microphones with the adjusted reception direction.
[0211] A method according to one embodiment may further include, when a plurality of speakers are identified based on the first image information, an operation of confirming the utterance order and the flow of the conversation content for the plurality of speakers, an operation of displaying virtual content individually for each of the plurality of speakers based on the utterance order and the flow of the conversation content, and an operation of applying a highlight effect to each of the virtual content.
[0212] The operation of displaying the virtual content according to one embodiment may include the operation of confirming the location of the first speaker based on the first image information, the operation of determining the display location of the virtual content based on the location of the first speaker, and the operation of displaying the virtual content according to the determined display location.
[0213] According to one embodiment, a non-transient computer-readable storage medium (or computer program product) storing one or more programs for performing a method of recognizing a speaker in an electronic device (101) may be described. According to one embodiment, the one or more programs may include instructions that, when executed by a processor (120) of the electronic device (101), acquire an audio signal through a plurality of microphones (310) while the electronic device is worn on a part of a user's body, acquire a first image information based on at least one camera (180) in response to the acquisition of the audio signal, extract a function corresponding to situation information based on the first image information and the audio signal corresponding to the conversation content of the first speaker included in the first image information, display virtual content corresponding to the function through a display (160) in response to the extraction of the function, and perform the function in response to an input to the virtual content.
[0214] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0215] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0216] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0217] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0218] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0219] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101) that can be worn on the head by a user, At least one camera (180); Housing; A display (160) supported by the above housing and outputting visual information; Multiple microphones (310); A communication module (190) including a communication circuit; A processor (120) including a processing circuit; and Includes memory (130) for storing instructions, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, With the above electronic device (101) worn on a part of the user's body, an audio signal is acquired through the plurality of microphones (310), and In response to the acquisition of the above audio signal, first image information is acquired based on the at least one camera (180), and Based on the first image information and the audio signal corresponding to the utterance content of the first speaker included in the first image information, a function corresponding to the situation information is extracted, and In response to the extraction of the above function, virtual content for executing the above function is displayed through the display (160), and An electronic device that executes the above function in response to input regarding the above virtual content.
2. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Based on the speech-related information (333) stored in the memory (130), the past conversation record of the first speaker and the function execution record mapped to the past conversation record are checked, and Based on the confirmed past conversation records and the confirmed function execution records, the function related to the first speaker is extracted, and An electronic device that updates the utterance-related information (333) corresponding to the first utterer when performing the above function.
3. In Paragraph 2, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Based on the confirmed past conversation record above, virtual content corresponding to the first speaker is displayed, and An electronic device that displays the identified past conversation record in response to input regarding the virtual content.
4. In Paragraph 3, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, In response to the conversation time with the first speaker exceeding a set threshold time, a first summary for the first speaker is generated based on the past conversation record for the first speaker, and An electronic device that displays the original text corresponding to the first summary in response to input for the first summary.
5. In Paragraph 1, A pupil tracking camera for detecting the user's pupil movement; further comprising, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Based on pupil movements identified using the pupil tracking camera, the user recognizes the first speaker they are gazing at, and An electronic device that displays a visual emphasis effect along the outer side of the recognized first speaker.
6. In Paragraph 5, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Adjusting the reception direction of the plurality of microphones (310) toward the recognized first speaker, and An electronic device for receiving audio information containing conversation content with the first speaker based on the plurality of microphones (310) having the receiving direction adjusted.
7. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, When multiple speakers are identified based on the first image information above, the utterance order and flow of conversation content for the multiple speakers are verified, and Based on the above utterance order and the flow of the above conversation content, virtual content is displayed individually for each of the plurality of speakers, and An electronic device that displays visual emphasis effects on virtual content corresponding to a speaker during speech.
8. In Paragraph 1, The above functions include at least one of the registration of a schedule, the transfer of a file, the search for a location, and the execution of a set application.
9. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Based on the first image information above, the location of the first speaker is confirmed, and Determining the display location of the virtual content based on the location of the first speaker, and An electronic device that displays the virtual content according to the above-determined display position.
10. A method for recognizing a speaker in an electronic device (101) that can be worn on the head by a user, The operation of acquiring an audio signal through a plurality of microphones (310) while the above electronic device (101) is worn on a part of the user's body; In response to the acquisition of the above audio signal, an operation of acquiring first image information based on at least one camera (180); An operation of extracting a function corresponding to situation information based on the audio signal corresponding to the first image information and the conversation content of the first speaker included in the first image information; In response to the extraction of the above function, an operation of displaying virtual content corresponding to the above function through the display (160); and A method comprising: an operation that performs the above function in response to input regarding the above virtual content.
11. In Paragraph 10, An operation to verify past conversation records and past function execution records for the first speaker based on speaker-related information stored in memory; An operation to extract the function corresponding to the first speaker based on the confirmed past conversation record and the confirmed past function execution record; and A method further comprising, when performing the above function, an operation of updating speaker-related information corresponding to the first speaker.
12. In Paragraph 11, The operation of generating a first summary for the first speaker based on the confirmed past conversation record; The operation of displaying virtual content corresponding to the first summary through the display; and A method further comprising: an action of displaying at least one of the first summary and the identified past conversation record in response to input regarding the virtual content.
13. In Paragraph 12, An operation to generate the first summary for the first speaker based on the past conversation record for the first speaker in response to the conversation time with the first speaker exceeding a set threshold time; and A method further comprising: an operation of displaying the original text corresponding to the first summary in response to input for the first summary.
14. In Paragraph 10, When multiple speakers are identified based on the first image information above, an operation to verify the speech order and flow of conversation content for the multiple speakers; An operation of individually displaying virtual content for each of the plurality of speakers based on the above utterance order and the flow of the above conversation content; and A method further comprising the action of applying a highlight effect to each of the above-mentioned virtual contents.
15. A non-transient computer-readable storage medium storing one or more programs for performing a method of recognizing a speaker in an electronic device (101), When one or more of the above programs are executed by the processor (120) of the electronic device (101), The operation of acquiring an audio signal through a plurality of microphones (310) while the above electronic device is worn on a part of the user's body; In response to the acquisition of the above audio signal, an operation of acquiring first image information based on at least one camera (180); An operation of extracting a function corresponding to situation information based on the audio signal corresponding to the first image information and the conversation content of the first speaker included in the first image information; In response to the extraction of the above function, an operation of displaying virtual content corresponding to the above function through the display (160); and A computer-readable storage medium comprising instructions for performing the operation of performing the above function in response to input for the above virtual content.