Electronic device for providing information about sound source, and operation method thereof
The HMD device addresses the lack of sound source description in wearable devices by detecting and describing sound sources using microphones and cameras, enhancing user experience and functionality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-21
AI Technical Summary
Existing wearable electronic devices, such as HMDs, lack effective methods to provide users with descriptive information about detected sound sources based on location and behavior information, limiting their functionality and user experience.
The HMD device includes a processor and memory to detect sound sources using microphones, acquire image information through cameras, identify sound source objects, and generate descriptive information using user-related data, providing notifications accordingly.
Enhances user experience by accurately identifying and describing sound sources, improving functionality and usability of wearable devices.
Smart Images

Figure KR2025018185_21052026_PF_FP_ABST
Abstract
Description
Electronic device providing information about a sound source and method of operation thereof
[0001] The present disclosure relates to an electronic device and a method of operation that provide information about a sound source.
[0002] With the advancement of digital technology, electronic devices are being provided in various forms, such as smartphones, tablet PCs, or PDAs. Electronic devices are also being developed in wearable forms to enhance portability and user accessibility.
[0003] Electronic devices developed in a form that users can wear are being developed in the form of wearable electronic devices such as AR glasses (augmented reality glasses), VST (video see-through) devices, and HMD (head-mounted display) devices to provide virtual spaces in virtual environments, and the various services and additional functions provided by wearable electronic devices are gradually increasing. To enhance the utility value of these electronic devices and satisfy the needs of diverse users, telecommunications service providers or electronic device manufacturers are competitively developing electronic devices to provide various functions and differentiate themselves from other companies. Accordingly, the various functions provided through wearable electronic devices are also becoming increasingly sophisticated. For example, wearable devices can support a function that provides information about audio sources to the user wearing the device.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] According to one embodiment, an HMD device may be provided. The HMD device may include at least one processor including a processing circuit and a memory including at least one storage medium for storing instructions. When the instructions are executed individually or collectively by the at least one processor, they may cause the HMD device to perform at least one operation. The at least one operation may include detecting a sound source from a sound signal obtained through one or more microphones, based on at least one of location information or behavior information of a user wearing the HMD device, to which descriptive information is provided (or meaningful sound information is provided). The at least one operation may include acquiring first image information of the sound source through a camera located in the direction of the sound source, based on direction information of the sound source. The at least one operation may include identifying whether a sound source object corresponding to the sound source is included within the image of the first image information. The at least one operation may include an operation of generating first descriptive information about the sound source using the first image information and user-related information based on identifying that the sound source object is included in the image. The at least one operation may include an operation of generating second image information including the sound source object based on identifying that the sound source object is not included in the image, and generating second descriptive information about the sound source using the second image information and user-related information. The at least one operation may include an operation of providing notification information including the first descriptive information or the second descriptive information.
[0006] According to one embodiment, a method of operation for an HMD device may be provided. The method of operation for an HMD device may include at least one operation. The at least one operation may include an operation of detecting a sound source for which descriptive information is provided (or meaningful sound information is provided) from a sound signal obtained through one or more microphones, based on at least one of location information or behavior information of a user wearing the HMD device. The at least one operation may include an operation of obtaining first image information for the sound source through a camera located in the direction of the sound source, based on direction information of the sound source. The at least one operation may include an operation of identifying whether a sound source object corresponding to the sound source is included within the image of the first image information. The at least one operation may include an operation of generating first descriptive information for the sound source using the first image information and user-related information, based on the identification that the sound source object is included within the image. The above at least one operation may include generating second image information containing the sound source object based on identifying that the sound source object is not included in the image, and generating second descriptive information about the sound source using the second image information and the user-related information. The above at least one operation may include providing notification information containing the first descriptive information or the second descriptive information.
[0007] According to one embodiment, a storage medium may be provided for storing at least one instruction readable by a computer. The at least one instruction may cause the HMD device to perform at least one operation when executed by at least a part of at least one processor of the HMD device. The at least one operation may include an operation of detecting a sound source for which descriptive information is provided (or meaningful sound information is provided) from a sound signal obtained through one or more microphones, based on at least one of location information or behavior information of a user wearing the HMD device. The at least one operation may include an operation of obtaining first image information for the sound source through a camera located in the direction of the sound source, based on direction information of the sound source. The at least one operation may include an operation of identifying whether a sound source object corresponding to the sound source is included within the image of the first image information. The at least one operation may include an operation of generating first descriptive information for the sound source using the first image information and user-related information based on the identification that the sound source object is included within the image. The above at least one operation may include generating second image information containing the sound source object based on identifying that the sound source object is not included in the image, and generating second descriptive information about the sound source using the second image information and the user-related information. The above at least one operation may include providing notification information containing the first descriptive information or the second descriptive information.
[0008] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0009] FIG. 1a is a block diagram of an electronic device in a network environment according to various embodiments of the present disclosure.
[0010] FIG. 1b is a drawing for illustrating a generative artificial intelligence system according to one embodiment of the present disclosure.
[0011] FIG. 2 is a drawing showing the configuration of a wearable electronic device according to one embodiment of the present disclosure.
[0012] FIGS. 3a to 3c are drawings showing the front and rear of a wearable electronic device according to one embodiment of the present disclosure.
[0013] FIG. 4 is a drawing of a wearable electronic device according to one embodiment of the present disclosure.
[0014] FIG. 5 is a flowchart illustrating a method for an HMD device to provide a notification based on a sound signal according to one embodiment of the present disclosure.
[0015] FIG. 6 is a diagram illustrating a method for an HMD device to provide a notification based on a sound signal according to one embodiment of the present disclosure.
[0016] FIG. 7 is a flowchart illustrating the operation of an HMD device detecting a sound source according to one embodiment of the present disclosure.
[0017] FIG. 8 is a flowchart illustrating the operation of an HMD device updating a reference sound signal for detecting a sound source according to one embodiment of the present disclosure.
[0018] FIG. 9a is a flowchart illustrating the operation of an HMD device acquiring image information about a sound source according to one embodiment of the present disclosure.
[0019] FIG. 9b is a flowchart illustrating the operation of an HMD device acquiring directional information about a sound source according to one embodiment of the present disclosure.
[0020] FIG. 10 is a signal flow diagram illustrating an operation in which an HMD device acquires image information about a sound source using a camera of a wearable electronic device connected to the HMD device, according to one embodiment of the present disclosure.
[0021] FIG. 11 is a flowchart illustrating the operation of an HMD device estimating a sound source object according to one embodiment of the present disclosure.
[0022] FIG. 12 is a diagram illustrating the operation of an HMD device generating image information about a sound source and descriptive information about a sound source using a generative AI model according to one embodiment of the present disclosure.
[0023] FIG. 13 is a diagram illustrating the operation of an HMD device updating notification information according to the user's situation, according to one embodiment of the present disclosure.
[0024] FIG. 14 is a diagram illustrating the operation of an HMD device updating notification information according to the user's situation, according to one embodiment of the present disclosure.
[0025] FIGS. 15 and 16 are drawings illustrating the operation of an HMD device adding new notification information according to the user's situation, according to one embodiment of the present disclosure.
[0026] FIG. 17 is a diagram illustrating the operation of an HMD device providing notification information about a sound source according to one embodiment of the present disclosure.
[0027] FIG. 18 is a diagram illustrating the operation of an HMD device providing notification information about a sound source according to one embodiment of the present disclosure.
[0028] FIG. 19 is a drawing illustrating the configuration of a first electronic device according to one embodiment of the present disclosure.
[0029] FIG. 20 is a drawing illustrating the configuration of a second electronic device according to one embodiment of the present disclosure.
[0030] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0031] FIG. 1a is a block diagram of an electronic device in a network environment according to various embodiments of the present disclosure.
[0032] Referring to FIG. 1a, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0033] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0034] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0035] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0036] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0037] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0038] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0039] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0040] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0041] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0042] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0043] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0044] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0045] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0046] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0047] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0048] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a communication module (192) (e.g., cellular communication module, short-range communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0049] The communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0050] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0051] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0052] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface) and exchange signals (e.g., commands or data) with each other.
[0053] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0054] The number of processors (120) may be one or more. According to one embodiment, the processors may be physically multiple and may have a structure that is physically contained within a single chip. For example, the processor (120) may have a structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0055] The processor (120) can control the operations of the electronic device (101) by executing instructions stored in the memory (130). For example, the processor (120) may correspond to a plurality of processors that divide and collectively perform a plurality of operations among the processors.
[0056] FIG. 1b is a drawing for illustrating a generative artificial intelligence system according to one embodiment of the present disclosure.
[0057] According to one embodiment, a user query / response interface (210b) may receive input (e.g., user input or data acquired or generated by an electronic device (e.g., the electronic device (101) of FIG. 1a). Data acquired or generated by the electronic device may include, for example, image or video data generated using a processor (e.g., the processor (120) of FIG. 1a), values received through a sensor (e.g., the sensor module (176) of FIG. 1a) or a sensor hub (e.g., external illumination, angle of the electronic device, temperature of a display (e.g., the display module (160) of FIG. 1a) or the electronic device, display size or expansion / reduction information, or images captured by an image sensor). User input may be in the form of natural language, touch coordinates or stylus coordinates acquired through a touch panel or digitizer included in the display, images and / or videos, but is not limited thereto. Additionally, context information may be transmitted along with the transmission of user input. Context information may include various at the time of user input Additional information may be included. For example, additional information may include information about the application currently being used by the user or the user's location information. Additionally, user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, user input may be in a non-natural language form, such as selecting a menu. The user query / response interface (210b) may output to the user the results of the generative artificial intelligence system (200b) and / or the results of analyzing the input. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user query / response interface (210b) may output to the user the results of the generative artificial intelligence system (200b).The output can be in the form of natural language or specific content, and it may also be provided in the form of an action requested by the user.
[0058] According to one embodiment, the AI framework (240b) can receive user input and coordinate and control each component necessary to perform the user's intent based on the user's query.
[0059] According to one embodiment, user input received from a user query / response interface (210b) may be transmitted to a prompt design component (241b). The prompt design component (241b) may be used to generate prompts suitable for inputting user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (241b) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (241b) may generate prompts by accessing a knowledge component containing user preference data, a prompt library, and prompt examples based on user input, and transmit the generated prompts to the LLM or LMM.
[0060] According to one embodiment, the API / Plug-in management component (242b) can perform the role of communicating with external information when there is a request for additional information when transmitting user input as input to a generative model. The API / Plug-in management component (242b) establishes a channel to communicate with the outside of the AI Interface via an API, and can enable access to various data sources (e.g., a knowledge repository (220b)) through the established channel. Additionally, if the API / Plug-in management component (242b) needs to perform an action that executes the user input as a final step rather than an intermediate result in an application or service, it can request such action from the application / service component (230b) via an API. Information obtained from the outside can be used to generate a prompt in the prompt design component (241b) along with the user input, or it can be transmitted as input to the generative model.
[0061] According to one embodiment, an output modification component (or refiner component) (243b) can fine-tune the output of a generative model. For example, the output modification component (243b) can verify whether the content generated through the LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the output modification component (243b) can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Furthermore, the output modification component (243b) can configure and provide hints to the user to avoid unwanted output.
[0062] According to one embodiment, a generative AI model (260b) may generally refer to an artificial intelligence neural network that generates new forms of data based on user input information. The generative AI model (260b) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, generative adversarial networks (GANs) and variational autoencoders (VAEs), and examples of diffusion-based generative models using VAEs and Transformer structures can be cited. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples of which include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also large multimodal models (LMMs) that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.
[0063] FIG. 2 is a drawing showing the configuration of a wearable electronic device according to one embodiment of the present disclosure.
[0064] Referring to FIG. 2, according to one embodiment, a wearable electronic device (200) (e.g., the electronic device (101) of FIG. 1a) may include a light output module (211), a display member (201), a camera module (250), and / or a speaker (261).
[0065] According to one embodiment, a light output module (211) (e.g., a display module (160) of FIG. 1a) may include a light source capable of outputting an image and a lens that guides the image to a display member (201). The light output module (211) may include, for example, a liquid crystal display, a digital mirror device, a liquid crystal on silicon display, an organic light emitting diode and / or a micro light emitting diode (micro LED).
[0066] According to one embodiment, a display member (201) (e.g., a display module (160) of FIG. 1a) may include an optical waveguide (e.g., a waveguide). According to one embodiment, an output image of an optical output module (211) incident on one end of the optical waveguide may propagate within the optical waveguide and be provided to a user. According to one embodiment, the optical waveguide may include at least one diffractive element (e.g., a diffractive optical element (DOE), a holographic optical element (HOE)) and / or a reflective element (e.g., a reflective mirror). For example, the optical waveguide may guide the output image of the optical output module (211) to the user's eye using at least one diffractive element or reflective element.
[0067] According to one embodiment, a camera module (250) (e.g., camera module (180) of FIG. 1a) can capture images (e.g., still images and / or video). According to one embodiment, the camera module (250) may be placed within a lens frame and around a display member (201). In the present disclosure, images may be interpreted to include video as well as still images.
[0068] According to one embodiment, the first camera module (251) can capture and / or recognize the trajectory of the user's eye (e.g., pupil, iris) or gaze. According to one embodiment, the first camera module (251) can periodically or non-periodically transmit information related to the trajectory of the user's eye or gaze (e.g., trajectory information) to a processor (e.g., processor (120) of FIG. 1a).
[0069] According to one embodiment, the second camera module (253) can capture an external image. For example, the second camera module (253) can capture an image of the external environment in the front direction of the wearable electronic device (200).
[0070] According to one embodiment, the third camera module (255) may be used for hand detection and tracking and user gesture (e.g., hand movements) recognition. According to one embodiment, the third camera module (255) may be used for 3 degrees of freedom (3DoF) and 6DoF head tracking, location (space, environment) recognition, and / or movement recognition. According to one embodiment, the second camera module (253) may be used for hand detection and tracking and user gesture recognition. According to one embodiment, at least one of the first camera module (251) to the third camera module (255) may be replaced with a sensor module (e.g., LiDAR sensor). For example, the sensor module may include at least one of a vertical cavity surface emitting laser (VCSEL), a diode, an infrared sensor, an infrared diode, and / or a photodiode.
[0071] According to one embodiment, a speaker (261) (e.g., the acoustic output module (155) of FIG. 1a) can output an acoustic signal (e.g., sound and / or virtual vibration sound). Although the speaker (261) has been described as being configured in a member that is mounted on the user's ear when the wearable electronic device (200) is worn in FIG. 2, it is not limited thereto and may be configured in other locations depending on the implementation of the wearable electronic device (200).
[0072] FIGS. 3a to 3c are drawings showing the front and rear of a wearable electronic device according to one embodiment of the present disclosure.
[0073] Referring to FIGS. 3a through 3c, according to one embodiment, at least one first camera module (311, 312) and at least one second camera module (313, 314, 315, 316), a depth sensor (317), and / or a second display (350) may be disposed on a first surface (310) of a housing for acquiring information related to the surrounding environment of a wearable electronic device (300) (e.g., the electronic device (101) of FIG. 1a).
[0074] According to one embodiment, at least one first camera module (311, 312) can capture an image of the outside of the wearable electronic device (300). For example, the first camera module (311, 312) can capture an image of the external environment in the front direction of the wearable electronic device (300).
[0075] According to one embodiment, at least one second camera module (313, 314, 315, 316) can acquire images while the wearable electronic device (300) is worn by a user. The second camera module (313, 314, 315, 316) may be used for hand detection, tracking, and user gesture (e.g., hand movements) recognition. The second camera module (313, 314, 315, 316) may be used for 3DoF, 6DoF head tracking, location (space, environment) recognition, and / or movement recognition. According to one embodiment, a first camera module (311, 312) may be used for hand detection and tracking and user gestures.
[0076] According to one embodiment, the depth sensor (317) may be configured to transmit a signal and receive a signal reflected from a subject, and may be used for purposes such as time of flight (TOF) to determine the distance to an object. Alternatively, or additionally, a second camera module (313, 314, 315, 316) may determine the distance to an object.
[0077] According to one embodiment, the second display (350) (and / or lens) may be placed on the first surface (310) of the wearable electronic device (300). According to one embodiment, the second display (350) may provide visual information to the outside of the wearable electronic device (300). For example, the second display (350) may be used to provide an alternative notification indicating the operating status of the first camera module (311, 312) in place of the light emitter (340).
[0078] According to one embodiment, camera modules (325, 326) for face recognition and / or a first display (321) (and / or a lens) may be disposed on the second surface (320) of the housing.
[0079] According to one embodiment, camera modules (325, 326) for face recognition adjacent to the first display (321) may be used to recognize the user's face or to recognize and / or track both of the user's eyes.
[0080] According to one embodiment, the first display (321) (and / or lens) may be disposed on the second surface (320) of the wearable electronic device (300). According to one embodiment, the wearable electronic device (300) may not include camera modules (315, 316) among a plurality of second camera modules (313, 314, 315, 316). Although not illustrated in FIG. 3a and 3b, the wearable electronic device (300) may further include at least one of the configurations illustrated in FIG. 2.
[0081] Referring to FIG. 3c, according to one embodiment, the wearable electronic device (300) may have a form factor (e.g., a head-mounted display (HMD)) for being worn on a user's head. The wearable electronic device (300) may further include a strap and / or a wearing member for being secured on a part of the user's body. The wearable electronic device (300) may include a volume button (331), a vent (333), a status indicator (335), and a power button (e.g., including a fingerprint recognition sensor) (337), and such configurations may be identically included in the wearable electronic device (300) illustrated in FIG. 3a and FIG. 3b. When worn on a user's head, it may provide a user experience based on augmented reality, virtual reality, and / or extended reality (or mixed reality). The wearable electronic device (300) configured in the form of an HMD may include configurations identical or similar to the components of FIG. 3a and FIG. 3b described above.
[0082] According to one embodiment, a speaker (318) (e.g., the acoustic output module (155) of FIG. 1a or the speaker (261) of FIG. 2) may output an acoustic signal (e.g., sound and / or virtual vibration sound). Although the speaker (318) has been described as being configured in a location adjacent to the vent (333) in FIG. 3a through 3c as an example, it is not limited thereto and may be configured in other locations depending on the implementation of the wearable electronic device (300).
[0083] FIG. 4 is a drawing of a wearable electronic device according to one embodiment of the present disclosure.
[0084] According to one embodiment, the wearable electronic device (400) may be, for example, earbuds configured as a pair and worn on each of the user's ears.
[0085] Referring to FIG. 4, according to one embodiment, a wearable electronic device (400) may include a housing (410). A space may be formed inside the housing (410). The housing (410) may include a first housing (411) and a second housing (412). The first housing (411) and the second housing (412) may be formed integrally. A microphone (440), a camera (450), and a speaker (460) may be placed inside the housing (410).
[0086] According to one embodiment, the wearable electronic device (400) may include a port (420). The port (420) may be coupled to a housing (310). The port (420) may protrude outward from the housing (410). The port (420) may be coupled to a second housing (412). Sound generated from a speaker (460) may be output to the outside of the housing (410) through the port (420).
[0087] According to one embodiment, the wearable electronic device (400) may include a grille (430). The grille (430) may be coupled to a housing (410). Sound from outside the housing (410) may pass through the grille (430) and enter into the housing (410). The housing (410) may include an opening, and the grille (430) may be placed in the opening.
[0088] According to one embodiment, the wearable electronic device (400) may include a microphone (440) (e.g., the input module (150) of FIG. 1A). The microphone (440) may be placed inside the housing (410). The microphone (450) may receive sound generated outside the housing (410). The microphone (450) may include a plurality of microphones (e.g., a first microphone (441) and a second microphone (442)).
[0089] According to one embodiment, the wearable electronic device (400) may include a camera (450) (e.g., the camera module (180) of FIG. 1a). The camera (450) can capture still images and videos within the field of view of the camera (450).
[0090] According to one embodiment, the wearable electronic device (400) may include a speaker (460) (e.g., the acoustic output module (155) of FIG. 1a). The speaker (460) may be placed inside a housing (410). The speaker (460) may output sound outside the housing (410).
[0091] FIG. 5 is a flowchart illustrating a method for an HMD device to provide a notification based on a sound signal according to one embodiment of the present disclosure.
[0092] Referring to FIG. 5, in operation 510, an HMD device (e.g., the electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) can detect a sound source that provides meaningful sound information from a sound signal acquired through at least one microphone. The sound source that provides meaningful sound information may be, for example, a sound source for which explanatory information or notifications need to be provided. An example of a sound source detection operation is described below with reference to FIG. 7.
[0093] According to one embodiment, an HMD device can detect (or identify) a sound source that provides meaningful sound information from a sound signal based on location information and / or activity information of a user wearing the HMD device. According to one embodiment, the location information of a user wearing the HMD device may be location information of the HMD device or information obtained based on the location information of the HMD device. Location information may be obtained, for example, using a positioning technology using GPS and / or communication (e.g., a positioning technology using Wi-Fi). Activity information may be obtained, for example, based on a sensor (e.g., an inertial sensor of the HMD device) and / or information of an application running on the HMD device. In the present disclosure, a sound source that provides meaningful sound information may be referred to as a meaningful sound source.
[0094] According to one embodiment, meaningful sound information may include, for example, sound-related information corresponding to the user's situation (e.g., sound-related information suitable for the user's situation). The user's situation may be associated with, for example, various factors associated with the user (e.g., the user's location and / or behavior). For example, the user's situation may be established based on the user's location information and behavior information. For example, various situations may be recognized depending on the user's location and behavior.
[0095] According to one embodiment, sound-related information corresponding to the user's situation may include, for example, a specific keyword (e.g., the user's name), a keyword corresponding to the user's situation (e.g., work-related keywords), and / or information about a scene or environment corresponding to the user's situation (e.g., a car horn sound, a home appliance notification sound). In the present disclosure, information about a scene or environment may be referred to as a scene token. Thus, by providing notifications based on sound sources that provide meaningful sound information, the HMD device can provide appropriate notifications that take the user's situation into account, unlike a method that simply recognizes and provides notifications based on a specific sound that is already set (e.g., a sound of calling the user's name).
[0096] According to one embodiment, the operation of detecting a sound source may include an operation of updating at least one meaningful keyword (reference keyword) and at least one meaningful scene token (reference scene token) corresponding to the user's situation based on at least one of the user's location information or behavior information. According to one embodiment, the operation of detecting a sound source may include an operation of performing keyword matching between at least one keyword and at least one reference keyword in response to at least one keyword being obtained from a sound signal, an operation of performing scene matching between at least one scene token and at least one reference scene token in response to at least one scene token being obtained from a sound signal, and / or an operation of determining whether a sound source providing meaningful sound information is detected based on the results of keyword matching and scene matching. An example of the operation of determining whether such a meaningful sound source is detected is described below with reference to FIG. 7.
[0097] According to one embodiment, at least one microphone may be included in an HMD device and / or an electronic device connected to the HMD device (e.g., the electronic device (102) of FIG. 1a) or a wearable electronic device (e.g., the wearable electronic device (400) of FIG. 4). For example, the HMD device may acquire a plurality of sound signals (e.g., a composite sound signal) by using microphones included in the HMD device, microphones included in the wearable electronic device, or a combination of microphones included in the HMD device and the wearable electronic device. Each of the plurality of sound signals acquired through these various combinations of microphones (microphone arrays) may include spatial information. The spatial information may include, for example, information regarding the location and / or orientation of the microphone acquiring the corresponding sound signal. In the present disclosure, the electronic device or wearable device connected to the HMD device may include an electronic device or wearable device capable of communicating with the HMD device.
[0098] According to one embodiment, in operation 520, the HMD device can acquire (e.g., capture) first image information about the sound source through a camera positioned in the direction of the sound source based on the direction information of the sound source. An example of the operation to acquire the first image information is described below with reference to FIGS. 9a to 10. The direction information of the sound source can be acquired, for example, through the performance of sound source localization. Sound source localization can be performed, for example, using a TDOA (time difference of arrival) based method and an AoA (angle of arrival) based method, which will be described below.
[0099] According to one embodiment, the HMD device can acquire (e.g., calculate) location information and / or direction information of a sound source based on Time of Arrival (TDOA) information indicating the time difference in which a sound signal of a sound source reaches each of a plurality of microphones. The plurality of microphones may be included in the HMD device and / or other wearable electronic devices connected to the HMD device. An example of such TDOA-based sound source localization is described below with reference to FIGS. 9a and 9b.
[0100] According to one embodiment, the HMD device can acquire (e.g., calculate) directional information of a sound source based on Angle of Arrival (AoA) information indicating the angle at which a sound signal from a sound source reaches a single microphone (e.g., a directional microphone). The single microphone may be included in the HMD device or in another wearable electronic device connected to the HMD device.
[0101] According to one embodiment, the HMD device can use direction information of a sound source to select a camera (hereinafter referred to as the first camera) located in the direction of the sound source among a plurality of cameras, activate the selected first camera, and acquire first image information regarding the sound source through the activated first camera. The first camera may be activated only for a specified period. The specified period may be set, for example, to a period sufficient to capture an image of the sound source through the first camera. In this way, to acquire image information regarding the sound source, by activating only the camera located in the direction of the sound source (e.g., temporarily activating it), battery consumption associated with camera activation can be reduced. Being located in the direction of the sound source may include, for example, that the sound source is located within the field of view of the camera.
[0102] According to one embodiment, the first camera may be included, for example, in an HMD device or a wearable electronic device connected to the HMD device (e.g., the wearable electronic device (400) of FIG. 4). If the camera located in the direction of the sound source is not included in the HMD device, the HMD device cannot obtain image information about the sound source through its own camera(s), so it may cooperate with another electronic device (e.g., the wearable device (400) of FIG. 4) to obtain image information about the sound source through the camera of said electronic device. For example, if the first camera is included in the wearable electronic device connected to the HMD device, the HMD device may transmit a command to activate the first camera to the wearable electronic device through a communication circuit and receive first image information captured through the first camera from the wearable device through the communication circuit. Cases where a camera located in the direction of the sound source is not included in the HMD device may include, for example, cases where the sound source is located outside the field of view of the camera(s) of the HMD device (e.g., cases where the HMD device includes only camera(s) that capture images in the front direction, and the sound source is located in the side or rear direction of the HMD device). An example of an operation to acquire first image information regarding a sound source through a camera of a wearable electronic device connected to the HMD device is described below with reference to FIG. 10.
[0103] According to one embodiment, in operation 530, the HMD device can identify whether a sound source object corresponding to a sound source is included within the image (e.g., a still image or a video) of the first image information. The sound source object corresponding to the sound source may be, for example, an object containing a sound source. For example, if sound is generated through the speaker of a telephone, the sound source may be the speaker, and the sound source object corresponding to the sound source may be the telephone. For example, if the sound is a human voice, the sound source may be the human throat, and the sound source object corresponding to the sound source may be the human. If it is identified that a sound source object is included within the image, operation 540 may be performed. If it is identified that a sound source object is not included within the image, operation 550 may be performed. In the present disclosure, the sound source object may be referred to as a sound source generating object.
[0104] According to one embodiment, in operation 540, based on the identification that a sound source object is included in the image, the HMD device may generate first descriptive information about the sound source using first image information and / or user-related information. The first descriptive information may include descriptive information associated with the sound source and / or sound source object. Thus, when a sound source object is included in the image captured through the first camera, descriptive information about the sound source may be generated using the captured image without generating additional separate image information.
[0105] According to one embodiment, in operation 550, based on identifying that a sound source object is not included in the image, the HMD device may generate second image information including the sound source object and generate second descriptive information about the sound source using the second image information and / or user-related information. The second descriptive information may include descriptive information associated with the sound source and / or sound source object. In this way, if a sound source object is not included in the image captured through the first camera, separate image information including the sound source object may be additionally generated, and descriptive information about the sound source may be generated using the generated image. In the present disclosure, user-related information may be referred to as user meta-information.
[0106] According to one embodiment, the second image information may be generated using a generative AI model (e.g., the generative AI model (260b) of FIG. 1b) (e.g., a large vision model (LVM)). The generative AI model may be configured to receive prompt data generated based on information about the sound source and user metadata as input data, and to output second image information including a sound source object. An example of an operation to generate second image information about a sound source using a generative AI model is described below with reference to FIG. 12.
[0107] According to one embodiment, information regarding a sound source may include information regarding an estimated sound source object and / or audio information regarding the sound source. The estimation of a sound source object is described below with reference to FIG. 11. When audio information regarding a sound source is input to a generative AI model as information regarding a sound source, the generative AI model may estimate a sound source object corresponding to the sound source based on audio information regarding the sound source and user metadata, and generate second image information including the estimated sound source object. When information regarding an estimated sound source object is input to a generative AI model as information regarding a sound source, the generative AI model may generate second image information including the estimated sound source object without estimating the sound source object, based on information regarding the estimated sound source object and user metadata.
[0108] According to one embodiment, user meta information may include information for identifying a user and / or additional information associated with the user. User meta information may include, for example, user identification information, user schedule information, user behavior information, user activity information, user interest information, user preference information and / or information regarding the usage history of applications used by the user (e.g., mail, messages, calls, other applications). By using such user meta information to generate second image information, the second image information can be generated as an appropriate image that allows the user to sufficiently understand the corresponding sound source.
[0109] According to one embodiment, in operation 560, the HMD device may provide notification information including first descriptive information or the second descriptive information. According to one embodiment, the HMD device may also provide a related image (e.g., first image information or second image information) along with the notification information.
[0110] According to one embodiment, the first explanatory information and the second explanatory information may include substantially the same explanatory information, but are not limited thereto. Thus, by providing notification information including the first explanatory information or the second explanatory information, the HMD device can provide specific information tailored to the user's situation regarding the provided notification, unlike a method of providing only a pre-set sound as a notification. Through this, the user receiving the notification can clearly understand the situation regarding the notification and take action. Furthermore, as described above, since the provided notification is a notification regarding a meaningful sound source, it may be a notification tailored to the user's situation. An example of a notification for providing the first explanatory information is described below with reference to FIG. 17, and an example of a notification for providing the second explanatory information is described below with reference to FIG. 18.
[0111] According to one embodiment, the first description information and the second setting information may be generated using a generative AI model (e.g., the generative AI model (260b) of FIG. 1b (e.g., LLM)). The generative AI model may be configured to generate the first description information by receiving prompt data generated based on information about the sound source, first video information, and / or user metadata as input data, or to generate the second description information by receiving prompt data generated based on information about the sound source, second video information, and user metadata as input data. By using such user metadata to generate the description information, the description information can be generated as appropriate description information that allows the user to sufficiently understand the corresponding sound source. An example of an operation to generate description information about a sound source using a generative AI model is described below with reference to FIG. 12.
[0112] According to one embodiment, a generative AI model used to generate descriptive information (e.g., first descriptive information and / or second descriptive information) may be the same as or different from a generative AI model used to generate second image information.
[0113] According to one embodiment, notification information may be provided visually or audibly. For example, notification information may be provided visually through a display of an HMD device. For example, notification information may be provided audibly through a speaker of an HMD device or a speaker of a wearable electronic device connected to an HMD device.
[0114] According to one embodiment, the HMD device may generate descriptive information for each sound source when a plurality of significant sound sources are detected together or simultaneously. Each descriptive information for each significant sound source may be generated, for example, through operations 510 to 560 according to FIG. 5. According to one embodiment, the HMD device may analyze each descriptive information for a plurality of sound sources and set a priority for the descriptive information or a notification providing the descriptive information. The HMD device may provide notification levels differently according to the set priority.
[0115] Meanwhile, if the direction of a detected significant sound source is located within a range where it cannot be captured through the camera of the HMD device or another electronic device connected to the HMD device—that is, if the sound source object cannot be captured through any camera—the HMD device may not perform the action of activating the camera. In this case, the HMD device can generate a virtual image containing the sound source object using a generative AI model without performing the action of capturing the sound source object through the camera.
[0116] Meanwhile, according to an embodiment, when sufficient power is supplied to, for example, an HMD device or a wearable electronic device connected to an HMD device, the HMD device may, in response to the detection of a significant sound source, activate all or multiple cameras that can be activated, rather than only activating the first camera located in the direction of the sound source, to obtain image information about the sound source.
[0117] FIG. 6 is a diagram illustrating a method for an HMD device to provide a notification based on a sound signal according to one embodiment of the present disclosure.
[0118] The method of providing a notification of the embodiment of FIG. 6 may be an example of the method of providing a notification of the embodiment of FIG. 5.
[0119] Referring to FIG. 6, in operation 610, an HMD device (e.g., electronic device (101) of FIG. 1a, wearable electronic device (200) of FIG. 2, or wearable electronic device (300) of FIG. 3a to 3c) may acquire at least one sound signal (e.g., audio signal) and perform time-frequency analysis on at least one sound signal. As described above, at least one sound signal may be acquired, for example, through at least one microphone of the HMD device and / or at least one microphone of a wearable electronic device (e.g., wearable electronic device (400) of FIG. 4) connected to the HMD device.
[0120] According to one embodiment, time-frequency analysis can be performed, for example, using the short-time Fourier transform (STFT) method. The STFT method may be a method for converting a sound signal from the time domain to the frequency domain by dividing a long signal (e.g., a sound signal) into short time intervals (windows) and performing a Fourier transform on each interval. Through the STFT method, patterns of change in the sound signal along the time and frequency axes can be identified, and consequently, a time-frequency spectrum (or spectrogram) can be obtained. The time-frequency spectrum can provide information on the magnitude and phase of the frequency components corresponding to each time interval. The time-frequency spectrogram thus obtained can be used to distinguish various sound components by including both frequency and time information.
[0121] According to one embodiment, an HMD device can separate a voice signal and a scene signal from at least one sound signal based on the results of time-frequency analysis. The voice signal may include, for example, a signal containing a sound corresponding to a human voice. Information associated with a human voice may be identified from such a voice signal. The scene signal may include, for example, a signal containing a sound that is not a human voice (e.g., ambient sound, background sound, environmental sound). Information associated with a scene or environment (e.g., scene token information) may be identified from such a scene signal. The HMD device may extract features for each of the separated voice signal and scene signal and apply a matching algorithm to detect a meaningful sound source.
[0122] According to one embodiment, in operation 620a, the HMD device can perform speech recognition on a speech signal. Speech recognition can be performed, for example, using automatic speech recognition (ASR) technology. ASR can be used, for example, for speech-to-text (STT). For example, the HMD device can recognize the speech of a speech signal using ASR technology and convert the speech into text data.
[0123] According to one embodiment, in operation 621a, the HMD device can extract keywords from text data. Keyword extraction may be performed using a specified keyword extraction algorithm. The specified keyword extraction algorithm may be, for example, a statistical analysis-based algorithm, a machine learning-based algorithm, a deep learning-based algorithm, or a rule-based algorithm, but is not limited thereto.
[0124] According to one embodiment, in operation 622a, the HMD device may perform a keyword match. For example, the HMD device may check whether the extracted keyword matches a meaningful keyword (reference keyword). The meaningful keyword may be selected (or updated) based on the user's location information and / or activity information. The meaningful keyword may be associated with, for example, the user's location information, activity information, and / or situation. The meaningful keyword may be a keyword corresponding to the user's situation set based on, for example, the user's location information and / or activity information (e.g., a context-appropriate keyword). An example of updating the meaningful keyword based on the user's location information and / or activity information is described below with reference to FIG. 8. In the present disclosure, a keyword that matches the meaningful keyword may be referred to as a matching keyword.
[0125] According to one embodiment, in operations 620b and 621b, the HMD device may perform a scene classification operation for a scene signal to generate at least one scene token. In operation 622b, the HMD device may perform a scene-match. For example, the HMD device may determine whether a scene token matches a meaningful scene token (reference scene token). A meaningful scene token may be selected (or updated) based on the user's location information and / or activity information. A meaningful scene token may be associated with, for example, the user's location information, activity information, and / or context. A meaningful scene token may be a scene token corresponding to the user's context set based on, for example, the user's location information and / or activity information (e.g., a context-appropriate scene token). An example of updating a meaningful scene token based on the user's location information and / or activity information is described below with reference to FIG. 8. In the present disclosure, a scene token that matches a meaningful scene token may be referred to as a matching scene token.
[0126] According to one embodiment, the HMD device may determine that a sound source (significant sound source) providing meaningful sound information from a sound signal is detected when at least one matching keyword and / or at least one matching scene token is identified. The HMD device may determine that a sound source (significant sound source) providing meaningful sound information from a sound signal is not detected when neither the matching keyword nor the matching scene token is identified.
[0127] The processing operations for the voice signals described above (e.g., operations 620a, 621a, 622a) and the processing operations for the scene signals (e.g., operations 620b, 621b, 622b) may be performed in parallel or simultaneously.
[0128] According to one embodiment, in operation 630, the HMD device may perform sound source localization. For example, the HMD device may perform sound source localization in response to the detection of a significant sound source. For a description of sound source localization, refer to, for example, the description of operation 520 of FIG. 5. Through such sound source localization, location information and / or direction information of a significant sound source may be obtained.
[0129] According to one embodiment, in operation 631, the HMD device can select and activate a camera located in the direction of the sound source based on the direction information of the sound source. For a description of the selection and activation of the camera, refer to, for example, the description of operation 520 of FIG. 5.
[0130] According to one embodiment, in operation 632, the HMD device can acquire first image information about a sound source through an activated camera. For the description of operation 632, for example, the description of operation 520 of FIG. 5 may be referenced.
[0131] According to one embodiment, in operation 640, the HMD device can identify whether a sound source object exists within the image of the first image information. For a description of operation 640, refer to, for example, the description of operation 530 of FIG. 5. If no sound source object exists within the image, operation 650 may be performed. If a sound source object exists within the image, operation 660 may be performed.
[0132] According to one embodiment, in operation 650, the HMD device may generate a prompt for generating second image information for a sound source based on user metadata. In operations 651 and 652, the HMD device may generate second image information for a sound source based on the generated prompt using a generative AI model (e.g., LVM). For a description of operations 650, 651, and 652, refer to, for example, the description of operation 550 of FIG. 5.
[0133] According to one embodiment, in operation 660, the HMD device may generate a prompt for generating first descriptive information (or second descriptive information) about a sound source based on information about the sound source, user metadata, and / or first visual information (or second visual information) about the sound source. In operations 661 and 662, the HMD device may generate first descriptive information (or second descriptive information) about the sound source based on the generated prompt using a generative AI model (e.g., LMM). For a description of operations 660, 661, and 662, refer to, for example, the description of operations 540 and 550 of FIG. 5.
[0134] According to one embodiment, in operation 670, the HMD device may provide notifications and information. For example, the HMD device may provide notification information including first explanatory information or second explanatory information. For a description of operation 670, for example, the description of operation 560 of FIG. 5 may be referenced.
[0135] FIG. 7 is a flowchart illustrating the operation of an HMD device detecting a sound source according to one embodiment of the present disclosure.
[0136] FIG. 8 is a flowchart illustrating the operation of an HMD device updating a reference sound signal for detecting a sound source according to one embodiment of the present disclosure.
[0137] The embodiments of FIGS. 7 and 8 may be, for example, examples of operation 510 of FIG. 5.
[0138] According to one embodiment, in operation 710, an HMD device (e.g., the electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) may update (or select) reference sound information including at least one meaningful keyword (reference keyword) and at least one meaningful scene token (reference scene token) based on the user's location information and / or activity information. In the present disclosure, reference sound information may be referred to as meaningful sound information.
[0139] According to one embodiment, a sound calling a user is information that must always be recognized and can always be included in reference sound information regardless of the user's location or activity. Additionally, the reference sound information may include at least one keyword (meaningful keyword) and / or at least one scene token (meaningful scene token) appropriate to the context based on the user's location and activity. Table 1 below illustrates an example of reference sound information based on the user's location information and activity information.
[0140] User Location (Location Information) User Activity (Activity Information) Reference Sound Information (Meaningful Sound Information) Working in the office - User Name - Work-related keywords Company premises Walking - User Name - Car horn sound Waiting at a bus stop - User Name - Subway, bus, etc. arrival notification sound Playing a game or watching a video at home - User Name - Doorbell sound - Appliance notification sound - Baby crying sound
[0141] Referring to Table 1, for example, in a situation where a user is wearing an HMD device and performing work in a company office, the reference sound information corresponding to that situation may include information on reference keywords such as the user's name and work-related keywords. For example, in a situation where a user is wearing an HMD device and walking at a company workplace, the reference sound information corresponding to that situation may include information on reference scene tokens such as the user's name and a vehicle horn sound. In this case, since the user is moving even though they are within the company, work-related keywords may not be meaningful sound information. For example, in a situation where a user is wearing an HMD device and waiting at a bus stop, the reference sound information corresponding to that situation may include information on reference scene tokens such as the user's name and notification sounds of transportation means such as subways and buses. For example, in a situation where a user is wearing an HMD device at home and playing games or watching videos, the reference sound information corresponding to that situation may include information on reference scene tokens such as the user's name and a doorbell sound, a home appliance notification sound, or a baby crying sound.
[0142] As such, reference sound information containing meaningful keywords and / or scene tokens can be configured differently depending on the situation corresponding to the user's location and behavior. Therefore, in order to detect meaningful sound sources appropriate to the user's situation, the HMD device needs to select or update appropriate reference sound information corresponding to the user's situation based on the user's location and behavior information.
[0143] Hereinafter, with reference to FIG. 8, an example of a method for updating (or selecting) reference sound information based on user location information and / or activity information is described.
[0144] According to one embodiment, in operation 810, the HMD device can determine whether the in / out status of a specific zone is recognized by checking whether the user has entered (in) or left (out) a specific zone based on the user's location information. The in / out status of a specific zone may be determined, for example, based on a geo-fence, but is not limited thereto. In operation 820, the HMD device can determine whether the user's activity is recognized based on the user's activity information. In operation 830, the HMD device can identify that the in / out status of a specific zone and / or the user's activity is recognized. If the in / out status of a specific zone is recognized, or the user's activity is recognized, or if both the in / out status of a specific zone and the user's activity are recognized, operation 840 may be performed. If neither the in / out status of a specific zone nor the user's activity is recognized, operation 810 may be performed again. In operation 840, the HMD device may update (or select) reference sound information including at least one meaningful keyword and / or at least one meaningful scene token based on the recognized information. Through this operation of updating reference sound information, reference sound information suitable for the situation corresponding to the user's location and activity may be updated.
[0145] The reference sound information selected (or updated) in this way can be used to determine whether keywords and / or scene tokens extracted from the sound signal correspond to meaningful keywords and / or scene tokens. For example, if a keyword extracted from the sound signal matches a keyword included in the updated reference sound information, that keyword can be identified as a meaningful keyword (matching keyword), and if a scene token extracted from the sound signal matches a scene token included in the updated reference sound information, that scene token can be identified as a meaningful scene token (matching scene token).
[0146] According to one embodiment, in operation 720, the HMD device can acquire a sound signal through at least one microphone. For a description of operation 720, refer to, for example, the description of operation 510 of FIG. 5 or operation 610 of FIG. 6.
[0147] According to one embodiment, in operation 730, the HMD device can identify whether at least one matching keyword and / or at least one matching scene token exists within the sound signal based on updated reference sound information. For a description of operation 730, refer to, for example, the description of operation 510 of FIG. 5 or operations 620a, 621a, 622a and operations 620b, 621b, 622b of FIG. 6. If at least one matching keyword, at least one matching scene token, or both at least one matching keyword and at least one matching scene token exist, operation 740 may be performed. If neither at least one matching keyword nor at least one matching scene token exists, operation 720 may be performed again.
[0148] According to one embodiment, in operation 740, the HMD device can identify that a sound source providing meaningful sound information (meaningful sound source) is detected when it is identified that at least one matching keyword, at least one matching scene token, or both at least one matching keyword and at least one matching scene token exist. For the description of operation 740, refer to, for example, the description of operation 510 of FIG. 5 or operations 620a, 621a, 622a and operations 620b, 621b, 622b of FIG. 6.
[0149] FIG. 9a is a flowchart illustrating the operation of an HMD device acquiring image information about a sound source according to one embodiment of the present disclosure.
[0150] FIG. 9b is a flowchart illustrating the operation of an HMD device acquiring directional information about a sound source according to one embodiment of the present disclosure.
[0151] The embodiments of FIG. 9a and 9b may be, for example, examples of operation 520 of FIG. 5. In the embodiment of FIG. 9b, for convenience of explanation, a plurality of sound signals used for sound source localization are described as, for example, a plurality of microphones (e.g., a microphone of an earbud worn on the left ear and a microphone of an earbud worn on the right ear) included in a wearable electronic device (e.g., a wearable electronic device of FIG. 4) connected to an HMD device (e.g., an electronic device (101) of FIG. 1a, a wearable electronic device (200) of FIG. 2, or a wearable electronic device (300) of FIG. 3a to 3c).
[0152] Referring to FIG. 9a, according to one embodiment, in operation 910, the HMD device can calculate Time of Arrival (TDOA) information indicating the time difference when a sound signal of a sound source reaches each of a plurality of microphones. In operation 910, the HMD device can obtain direction information of the sound source based on the TDOA information. Below, with reference to FIG. 9b, an example of an operation to obtain direction information of the sound source based on the TDOA information will be described.
[0153] Referring to FIG. 9b, for example, a sound signal (901) of a sound source (e.g., human voice) may reach each microphone of a wearable electronic device (e.g., earbuds, headphones) worn on both ears of the user. The waveform of the voice sound signal (901) may be the same as waveform (960), the waveform of the sound signal received by the microphone of the wearable electronic device worn on the left ear may be the same as the first waveform (961), and the waveform of the sound signal received by the microphone of the wearable electronic device worn on the right ear may be the same as the second waveform (962). At this time, due to the difference in distance from the sound source, the first waveform (961) may have a time delay compared to the second waveform (962). The HMD device analyzes the first waveform (961) and the second waveform (962) to determine the difference in arrival time ( Can calculate ) and the difference in arrival time ( The distance (D) and / or direction (α) between the user and the sound source can be calculated (or estimated) based on the distance between the user and the sound source and / or the distance between the two microphones (e.g., 2r). For example, to calculate the distance (D) and / or direction (α) between the user and the sound source, triangulation or beamforming methods may be used.
[0154] According to one embodiment, the HMD device can calculate (or estimate) the location of the sound source (origin location) by combining acquired directional information of the sound source. To estimate the location of the sound source, optimization techniques such as maximum likelihood estimation (MLE) or parameter estimation may be used. Through this, the location can be estimated more precisely.
[0155] According to one embodiment, the HMD device can correct the estimated position. To correct the estimated position, a filtering technique, such as a Kalman filter, may be used. Through this, errors in position estimation caused by external factors, such as noise, can be corrected.
[0156] According to one embodiment, in operation 930, the HMD device can select a first camera located in the direction of the sound source among a plurality of cameras using direction information of the sound source. In operation 940, the HMD device can activate the selected first camera. In operation 950, the HMD device can obtain first image information about the sound source through the activated first camera. For the description of operations 930 to 950, reference may be made to, for example, the description of 520 in FIG. 5 and 631 and 632 in FIG. 6.
[0157] FIG. 10 is a signal flow diagram illustrating an operation in which an HMD device acquires image information about a sound source using a camera of a wearable electronic device connected to the HMD device, according to one embodiment of the present disclosure.
[0158] The embodiment of FIG. 10 may be an example of operation 520 of FIG. 5.
[0159] As described above, for example, if a camera located in the direction of a sound source is not included in the HMD device (e.g., the wearable electronic device (101) of FIG. 1, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c), image information regarding the sound source cannot be obtained through the camera(s) of the HMD device. In this case, the HMD device can cooperate with another electronic device connected to the HMD device (e.g., the wearable device (400) of FIG. 4), as exemplified in FIG. 10, to obtain image information regarding the sound source through the camera of said electronic device. In the embodiment of FIG. 10, for convenience of explanation, the other electronic device cooperating with the HMD device is described as a wearable electronic device such as an earbud. However, the embodiment is not limited thereto. For example, cameras of various types of electronic devices (e.g., CCTV, smartphone) placed in the same space or an adjacent space as the HMD device can be used to acquire video information about the sound source.
[0160] Referring to FIG. 10, according to one embodiment, in operation 1010, the HMD device can select a first camera located in the direction of the sound source among a plurality of cameras using direction information of the sound source.
[0161] According to one embodiment, when an HMD device establishes a communication connection with another electronic device, it may receive and store location information, orientation information, and / or performance information of the other electronic device. Performance information may include, for example, information related to the camera of the electronic device, and the camera-related information may include, for example, the number of cameras, the field of view range of the cameras, and information regarding the shooting direction of the cameras, but is not limited thereto. Through this, the HMD device may know in advance where, in what direction, and within what field of view range the camera of the electronic device can capture images. Based on this, the HMD device may select a first camera of the other electronic device that is positioned in the direction of the sound source and capable of capturing images of the sound source.
[0162] According to one embodiment, in operation 1020, the HMD device can identify whether the first camera corresponds to a camera included in the HMD device. For example, the HMD device can identify whether the first camera corresponds to a camera included in the HMD device based on information regarding the field of view of a camera included in the HMD device. If the first camera is a camera included in the HMD device, operation 1030 may be performed. If the first camera is not a camera included in the HMD device, operation 1040 may be performed.
[0163] According to one embodiment, in operation 1030, the HMD device can activate its first camera and obtain first image information about a sound source captured through the activated first camera.
[0164] According to one embodiment, in operation 1040, the HMD device may transmit a command to activate the first camera of the wearable electronic device connected to the HMD device (or a wearable electronic device capable of communicating with the HMD device). The command may include, for example, information indicating the activation of the first camera, and / or information indicating the activation time of the first camera. In operation 1050, the wearable electronic device may, in response to receiving the command, activate its first camera and acquire (e.g., capture) first image information about a sound source through the activated first camera. The first camera may be activated for, for example, a specified period (e.g., a period indicated by the information indicating the activation time of the first camera). In operation 1051, the wearable electronic device may transmit the first image information acquired through the first camera of the wearable electronic device to the HMD device.
[0165] FIG. 11 is a flowchart illustrating the operation of an HMD device estimating a sound source object according to one embodiment of the present disclosure.
[0166] The embodiment of FIG. 11 may be an example of operation 550 of FIG. 5.
[0167] As described above, the image of a sound source object corresponding to the sound source may not be included in the image of the first image information obtained through a camera positioned in the direction of the sound source. For example, the image of the sound source object (e.g., intercom or doorbell) may not be included in the captured image because it is obscured by a wall. Even in this case, the HMD device (e.g., the electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) needs to generate second image information including the sound source object to generate descriptive information about the sound source, and the sound source object needs to be accurately estimated to generate the second image information including the sound source object.
[0168] According to one embodiment, in operation 1110, the HMD device can acquire location information of the user. In operation 1120, the HMD information can recognize the zone where the user is located based on the location information. Here, the zone may be a certain range of areas where the user mainly stays, such as a home or a company. In operation 1130, the HMD device can update (or select) map information for the recognized zone and / or a list of objects registered within the map of the map information. Here, the list of objects may be, for example, a list of objects that generate sound. For example, if the user is located at home, the HMD device can update map information of the home (e.g., room, living room layout) and / or a list of objects registered within the map of the home (e.g., home appliances such as an intercom, rice cooker, and washing machine placed in the home).
[0169] According to one embodiment, in operation 1140, the HMD device can identify whether a specific sound source (e.g., a significant sound source) is detected. For a description of operation 1140, refer to, for example, the description of operation 510 of FIG. 5. If a specific sound source is detected, operation 1150 may be performed. If a specific sound source is not detected, operation 1110 may be performed again.
[0170] According to one embodiment, in operation 1150, the HMD device may select candidate object information from an object list based on the detection of a specific sound source, based on a user location based on map information (map-based user location) and heading information. The candidate object information may include, for example, at least one candidate object selected from the object list. In response to the detection of a specific sound source, the HMD device may acquire user location information on a map based on map information and acquire heading information that provides information about the direction the user is looking based on a sensor of the HMD device (e.g., an inertial sensor). Based on the location information and heading information thus acquired, candidate object information corresponding to the user's location and direction may be selected. In the present disclosure, heading information may be referred to as direction information.
[0171] According to one embodiment, in operation 1160, the HMD device can analyze the correlation between candidate object information and a sound signal. The correlation analysis can be performed using correlation analysis techniques, such as the Pearson correlation coefficient and mutual information. In operation 1170, the HMD device can estimate a sound source object based on the results of the correlation analysis. For example, based on the results of the correlation analysis, the HMD device can select the object with the highest correlation among the candidate object(s) of the candidate object information as the sound source object. Through this process, even if a sound source object corresponding to the sound source cannot be identified using the image acquired through the camera, the sound source object corresponding to the sound source can be accurately estimated. The information regarding the sound source object estimated in this way can be used to generate image information about the sound source, for example, through a generative AI model (e.g., the generative AI model (1210) of FIG. 12).
[0172] FIG. 12 is a diagram illustrating the operation of an HMD device generating image information about a sound source and descriptive information about a sound source using a generative AI model according to one embodiment of the present disclosure.
[0173] The embodiment of FIG. 12 may be an example of operations 540 and 550 of FIG. 5.
[0174] Referring to FIG. 12, according to one embodiment, an HMD device (e.g., the wearable electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) can generate image information and / or descriptive information about a sound source using a generative AI model.
[0175] According to one embodiment, a first generative AI model (1210) (e.g., LVM) may be configured to receive prompt data generated based on information about a sound source and / or user metadata as input data, and to output second image information including a sound source object. LVM is a generative AI model for generating images, and may be, for example, a generative adversarial network (GAN), a variational auto encoder (VAE), or a diffusion-based generative model using a VAE and a transformer structure.
[0176] According to one embodiment, information about a sound source may include information about an estimated sound source object (e.g., a sound source object estimated through the operation of FIG. 11) and / or audio information about the sound source. According to one embodiment, user meta information may include information for identifying a user and / or additional information associated with the user. User meta information may include, for example, user identification information, user schedule information, user behavior information, user activity information, user interest information, user preference information and / or information about the usage history of applications used by the user (e.g., mail, message, call, other applications).
[0177] According to one embodiment, a second generative AI model (1220) (e.g., LLM) can generate descriptive information about a sound source by receiving prompt data generated based on information about a sound source, video information about a sound source, and / or user metadata as input data. For example, the second generative AI model (1220) can generate first descriptive information about a sound source by receiving prompt data generated based on information about a sound source, first video information acquired through a camera, and user metadata as input data. For example, the second generative AI model (1220) can generate second descriptive information about a sound source by receiving prompt data generated based on information about a sound source, second video information generated through the first generative AI model (1210), and user metadata as input data. The LLM is a generative AI model that generates language, and may be a model such as CHAT GPT-3 or CHAT GPT-4.
[0178] According to one embodiment, the first generative AI model (1210) may be the same as or different from the second generative AI model (1220).
[0179] According to one embodiment, the first generative AI model (1210) and / or the second generative AI model (1220) may be included in a server (e.g., the server (108) of FIG. 1a) or included in an HMD device. That is, the first generative AI model (1210) and / or the second generative AI model (1220) may be implemented to operate on a server or implemented as an on-device model on an HMD device.
[0180] Meanwhile, as described above, the information to be provided as a notification may vary depending on the current user's situation, such as the user's location and / or behavior. Accordingly, the user can determine whether information regarding keywords or scene tokens detected from sounds around the user corresponds to information that is meaningful to the current user (e.g., meaningful notification information or meaningful sound information for providing meaningful notification information). To this end, the user needs to identify and update meaningful sound information to receive notifications based on the user's location and / or behavior. Below, with reference to FIGS. 13 to 15, various examples of user interfaces (or screens) provided for the user to identify and update meaningful notification information (or meaningful sound information for providing meaningful notification information) will be described.
[0181] FIG. 13 is a diagram illustrating the operation of an HMD device updating notification information according to the user's situation, according to one embodiment of the present disclosure.
[0182] Referring to FIG. 13, an HMD device (e.g., the wearable electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) may display a first screen (1310) for updating meaningful notification information through a display. The update of meaningful notification information may include, for example, an update of meaningful sound information (reference sound information) of Table 1. For a description of the update of reference sound information, refer to, for example, the content described in FIG. 8.
[0183] According to one embodiment, the first screen (1310) may include information (1311) (e.g., an indicator) that displays the user's location (e.g., office) and / or a user interface part (1312) for updating meaningful notification information. The user interface part (1312) may include at least one selectable button (e.g., a button to check notification information, a button to cancel notification) along with guidance information related to the user's situation (e.g., "Mr. / Ms. OOO, you have started work. Updating key information to receive notifications during work."). The guidance information may be generated, for example, based on the user's location and / or behavior. As illustrated in FIG. 13, when the button to check notification information is selected based on user input (e.g., touch input, gesture input), meaningful notification information is updated, and the first screen (1310) may be switched to a second screen (1320). If the Cancel notification button is selected based on user input (e.g., touch input, gesture input), the HMD device may not provide the notification.
[0184] According to one embodiment, the second screen (1320) may include information (1321) that displays meaningful notification information in the current situation (e.g., a situation corresponding to the user's current location and actions). The information (1321) may include, for example, a username (e.g., "Hey OOO") and work-related keywords (e.g., "Pedometer"), as exemplified in Table 1, when the user is currently working in an office. Through this, the user can check information that allows them to receive meaningful notifications in their current situation. Thus, while the user is wearing an HMD device and performing work in an office, a username to detect when the user is called and keywords related to the user's assigned work may be provided as meaningful notification information. Keywords related to the user's assigned work may be automatically extracted and registered, for example, based on user meta-information.
[0185] FIG. 14 is a diagram illustrating the operation of an HMD device updating notification information according to the user's situation, according to one embodiment of the present disclosure.
[0186] Referring to FIG. 14, an HMD device (e.g., the wearable electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) may display a first screen (1410) for updating meaningful notification information through a display. The update of meaningful notification information may include, for example, an update of meaningful sound information (reference sound information) of Table 1. For a description of the update of reference sound information, refer to, for example, the content described in FIG. 8.
[0187] According to one embodiment, the first screen (1410) may include information (1411) (e.g., an indicator) that displays the user's location (e.g., home) and / or a user interface part (1412) for updating meaningful notification information. The user interface part (1412) may include at least one selectable button (e.g., a button to confirm notification information, a button to cancel notification) along with guidance information related to the user's situation (e.g., "Mr. / Ms. OOO, you have started a game. We are updating key information to be notified during the game."). The guidance information may be generated, for example, based on the user's location and / or actions. As illustrated in FIG. 14, when the button to confirm notification information is selected based on user input (e.g., touch input, gesture input), meaningful notification information is updated, and the first screen (1410) may be switched to a second screen (1420). If the Cancel notification button is selected based on user input (e.g., touch input, gesture input), the HMD device may not provide the notification.
[0188] According to one embodiment, the second screen (1420) may include information (1421) that displays meaningful notification information in the current situation (e.g., a situation corresponding to the user's current location and actions). The information (1421) may include, for example, a username (e.g., "Hey OOO") and a scene token associated with the current situation (e.g., "doorbell sound"), as exemplified in Table 1, when the user is currently playing a game at home. The scene token associated with the current situation may be automatically extracted and registered based on user meta-information, for example. Through this, the user can check information that allows them to receive meaningful notifications in their current situation. Thus, while the user is playing a game at home while wearing an HMD device, a username to detect when the user is called and a doorbell sound to identify a visitor while the user is playing a game may be provided as meaningful notification information.
[0189] FIGS. 15 and 16 are drawings illustrating the operation of an HMD device adding new notification information according to the user's situation, according to one embodiment of the present disclosure.
[0190] According to one embodiment, the HMD device may extract information regarding meaningful keywords (reference keywords) and / or meaningful scene talk (reference scene tokens) appropriate to the user's situation based on user meta-information and automatically register them. Additionally, depending on the case, the HMD device may manually (or directly) add, modify, or delete information regarding meaningful keywords and / or meaningful scene talk based on user input. For example, in a situation involving childcare at home, the HMD device may automatically register a baby crying sound as a scene token to receive notifications; however, it cannot provide notifications if the baby makes whining sounds instead of crying. Nevertheless, even in such cases, it is necessary to provide the sound to the user of the HMD device as a notification. Therefore, the HMD device needs to allow the user to add (or register) the sound as new notification information. Below, with reference to FIGS. 15 and 16, an example of a method for manually updating new notification information is described.
[0191] Referring to FIG. 15, an HMD device (e.g., the wearable electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) can display a first screen (1510) for adding (e.g., customizing) new notification information (e.g., meaningful notification information) through a display.
[0192] According to one embodiment, a first screen (1510) may include information (1511) (e.g., an indicator) for displaying the user's location (e.g., home) and / or a user interface portion (1512) for updating new notification information. The user interface portion (1512) may include at least one selectable button (e.g., a button to add new notification information, a button to modify existing notification information, a button to delete existing notification information) along with guidance information for the interface (e.g., "Customize notification information"). As illustrated in FIG. 15, when the button to add new notification information is selected according to user input (e.g., touch input, gesture input), the first screen (1510) may be switched to a second screen (1520). When the button to modify existing notification information or the button to delete existing notification information is selected according to user input (e.g., touch input, gesture input), the HMD device may perform at least one action to modify or delete existing notification information.
[0193] According to one embodiment, the second screen (1520) may include a user interface portion (1522) for adding new notification information. The user interface portion (1522) may include at least one selectable button (e.g., a start recording button, a cancel button) along with guidance information for the interface (e.g., "Recording a sound to be added as a new notification. When the recording is complete, the AI analyzes it and adds it to the notification information."). As illustrated in FIG. 15, when the start recording button is selected according to user input (e.g., touch input, gesture input), the HMD device can activate the recording function and record a sound to be added as a new notification. The HMD device can register the notification information (or sound information) for the recorded sound as meaningful notification information using an AI model (e.g., a generative AI model). The notification information thus registered may be added as notification information to the area where the user is currently located, for example, but is not limited thereto. When the registration of new notification information is completed, the second screen (1520) can be switched to a third screen (e.g., the screen (1610) of FIG. 16).
[0194] Referring to FIG. 16, the screen (1610) may include a user interface portion (1611) for setting newly added notification information as notification information for a specific activity only. The user interface portion (1611) may include at least one selectable button (e.g., end button, next button) along with guidance information for the interface (e.g., "Notification information to be received at OOO's house (baby whining sound) has been added. If you want to receive notifications only for a specific activity, please press "Next").
[0195] According to one embodiment, when the next button is selected based on user input (e.g., touch input, gesture input), the added notification information may be registered only as notification information for a specific activity (e.g., game activity). According to one embodiment, when the end button is selected based on user input (e.g., touch input, gesture input), the added notification information may be registered as notification information not limited to a specific activity. For example, the notification information may be registered as notification information for all types of activities in the area where the user is currently located.
[0196] Hereinafter, with reference to FIGS. 17 and 18, a method for providing notification information when a significant sound source is detected will be described.
[0197] FIG. 17 is a diagram illustrating the operation of an HMD device providing notification information about a sound source according to one embodiment of the present disclosure.
[0198] The embodiment of FIG. 17 may be an example of a scenario in which a user is performing work while wearing an HMD device (e.g., the wearable electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) in a company office, and a notification is provided regarding a situation in which a colleague calls the user. This scenario assumes, for example, a situation in which the user has a meeting to attend with a colleague but is late for the meeting, and the colleague belatedly realizes that there was a meeting, calls the user, and then immediately moves to the meeting room.
[0199] Referring to FIG. 17, the HMD device can detect a significant sound source based on a sound signal through at least one microphone. For example, the HMD device can identify that a significant sound source has been detected by confirming that a keyword obtained from the sound signal matches a user name. The HMD device can activate a first camera located in the direction of the identified sound source to obtain first image information about the sound source through the first camera. For example, as shown in FIG. 17, the first camera may be a first camera of a wearable electronic device (e.g., the wearable electronic device (400) of FIG. 4) connected to the HMD device (e.g., the camera of the right earbud). The HMD device can identify that a sound source object (e.g., a colleague) corresponding to the sound source is included in the image of the first image information, and can generate descriptive information about the sound source through a generative AI model based on the first image information and user meta information. The HMD device can provide notification information.
[0200] According to one embodiment, as illustrated in FIG. 17, notification information may be provided through the display of an HMD device. For example, the HMD device may display through the display a screen (1710) that shows a first information (1711) (e.g., an indicator) indicating the occurrence of notification information, along with a task currently being performed through the HMD device (e.g., a work task). Along with the display of the first information (1711), the HMD device may output a sound (e.g., a ding-dong) indicating the occurrence of notification information. As illustrated in FIG. 17, when the first information (1711) is selected according to user input (e.g., touch input, gesture input), a second information (1720) including descriptive information about the sound source may be displayed through the display. As illustrated, the second information (1720) may predict a situation in which a speech object calls the user and provide specific descriptive information thereabouts. For example, the second information (1720) may indicate that a colleague has called the user to attend a meeting and may provide information about the meeting (e.g., time, place, content). Through this, the user can understand the notification appropriately in the context. However, unlike the present disclosure, if the HMD device simply detects a specific preset sound and provides a notification, in a situation like the scenario of FIG. 17, the user can only confirm through the notification that a colleague has called them, but cannot understand why the colleague called the user when the colleague has already left.
[0201] FIG. 18 is a diagram illustrating the operation of an HMD device providing notification information about a significant sound source according to one embodiment of the present disclosure.
[0202] The embodiment of FIG. 18 may be an example of a scenario in which a user, for instance, wears an HMD device (e.g., the wearable electronic device (101) of FIG. 1a, the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3c) at home and plays a game, and provides a notification regarding a situation where a doorbell sound is detected. This scenario assumes, for instance, a situation where a credit card applied for by the user a few days ago has been delivered, and because the recipient must personally receive the mail due to the nature of registered mail, the user must stop playing the game and receive the mail before the mail carrier leaves. Meanwhile, in the embodiment of FIG. 18, unlike the embodiment of FIG. 17, the situation may be one in which a sound source object corresponding to a significant sound source is not included in the video captured through the camera. Therefore, the HMD device needs to generate video information containing the sound source object. The video information generated in this way can be used to generate descriptive information about the sound source.
[0203] Referring to FIG. 18, the HMD device can detect a significant sound source based on a sound signal through at least one microphone. For example, the HMD device can identify that a significant sound source has been detected by confirming that a scene token obtained from the sound signal matches a doorbell sound. The HMD device can activate a first camera located in the direction of the identified sound source to obtain first image information about the sound source through the first camera. For example, as shown in FIG. 18, the first camera may be a first camera of a wearable electronic device (e.g., the wearable electronic device (400) of FIG. 4) connected to the HMD device (e.g., the camera of the right earbud). The HMD device can identify (1801) that the image of the first image information does not contain a sound source object (e.g., a doorbell) corresponding to the sound source, and can estimate the sound source object (1802). As exemplified in FIG. 18, the sound source object may not be included in the video captured by the first camera due to the wall of the room. The HMD device can generate second video information containing the sound source object using a generative AI model based on information about the estimated sound source object and user metadata, and can generate descriptive information about the sound source using the generative AI model based on the second video information and user metadata (1803). The HMD device can provide notification information.
[0204] According to one embodiment, as illustrated in FIG. 18, notification information may be provided through the display of an HMD device. For example, the HMD device may display through the display a screen (1810) that shows a first information (1811) (e.g., an indicator) indicating the occurrence of notification information, along with a task currently being performed through the HMD device (e.g., a game task). Along with the display of the first information (1811), the HMD device may output a sound (e.g., a ding-dong) indicating the occurrence of notification information. As illustrated in FIG. 18, when the first information (1811) is selected according to user input (e.g., touch input, gesture input), a second information (1820) including descriptive information about the sound source may be displayed through the display. As illustrated, the second information (1820) may predict a situation where a doorbell is pressed and provide specific descriptive information about it. For example, the second piece of information (1820) may inform the user that a newly issued credit card has arrived by registered mail. Through this, the user can understand the notification appropriately in the context.
[0205] Meanwhile, according to one embodiment, the HMD device can generate descriptive information for each sound source when a plurality of significant sound sources are detected together or simultaneously. Each descriptive information for each significant sound source can be generated, for example, through operations 510 to 560 of FIG. 5. According to one embodiment, the HMD device can analyze each descriptive information for a plurality of sound sources to set a priority for the descriptive information or for a notification providing the descriptive information. The HMD device can provide notification levels differently according to the set priority. For example, when a user is playing a game while wearing the HMD device at home, if a doorbell sound due to registered mail delivery is detected together as in FIG. 18 and a sound of another person calling as in FIG. 17, the HMD device can analyze the descriptive information for each sound source using, for example, a generative AI model, and determine that receiving the mail is slightly more urgent. In this case, a notification by a doorbell sound, such as the information (1820) in Fig. 18, can be provided as a main notification, and a sound of another person calling, such as the information (1720) in Fig. 17, can be provided as a sub-notification. In this case, when the user selects a notification indicator (e.g., the information (1711) in Fig. 17 or the information (1811) in Fig. 18), information about the main notification is displayed first, and then information about the second notification can be provided when the next button is clicked.
[0206] FIG. 19 is a drawing illustrating the configuration of a first electronic device according to one embodiment of the present disclosure.
[0207] According to one embodiment, the first electronic device (1900) may be, for example, the electronic device (101) of FIG. 1a or the wearable electronic device (200) of FIG. 2, or the wearable electronic device (300) of FIG. 3a to 3b.
[0208] Referring to FIG. 19, according to one embodiment, the first electronic device (1900) may include, but is not limited to, at least one camera (1910), at least one microphone (1920), at least one speaker (1930), at least one sensor (1940), at least one display (1950), at least one communication circuit (1960), at least one processor (1970) and / or at least one memory (1980). The first electronic device (1900) may omit some of the components described above and may further include other components necessary to perform at least one operation.
[0209] According to one embodiment, at least one processor (1970) (e.g., processor (120) of FIG. 1a) may be electrically and / or operationally connected to other components (e.g., memory (1980)) by a configuration such as a communication bus.
[0210] According to one embodiment, at least one camera (1910) (e.g., camera module (180) of FIG. 1a) can capture still images and video. According to one embodiment, at least one camera (1910) may include one or more lenses, image sensors, image signal processors, or flashes.
[0211] According to one embodiment, at least one microphone (1920) (e.g., input module (150) of FIG. 1a) can receive a sound signal (or audio signal) containing commands or data to be used in a component (e.g., processor (1970)) of the first electronic device (1900) from outside the first electronic device (1900).
[0212] According to one embodiment, at least one speaker (1930) (e.g., the acoustic output module (155) of FIG. 1a) can output an acoustic signal to the outside of the first electronic device (1900).
[0213] According to one embodiment, at least one sensor (1940) (e.g., sensor module (176) of FIG. 1a) can detect the operating state of the first electronic device (1900) (e.g., power or temperature) or an external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, at least one sensor (1940) may include, for example, an inertial sensor, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0214] According to one embodiment, at least one display (1950) (e.g., the display module (160) of FIG. 1a) can provide information visually.
[0215] According to one embodiment, at least one communication circuit (1960) (e.g., communication module (190) of FIG. 1a) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between a first electronic device (1900) and an external electronic device, and the performance of communication through the established communication channel.
[0216] According to one embodiment, at least one processor (1970) can control the overall operation of the first electronic device (1900) and can perform at least one operation of the first electronic device (1900) (e.g., at least one of the operations described above in FIG. 1a through 18). The processor (1970) can perform operations or data processing regarding the control and / or communication of at least one other component of the first electronic device (1900). The processor (1970) may include at least one processing circuit that executes instructions stored in memory (1980).
[0217] According to one embodiment, at least one processor (1970) may include various processing circuits and / or multiple processors. One or more of the at least one processor (1970) may be configured to perform the various functions described in this disclosure individually and / or collectively. Where in this disclosure, "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms may cover, for example, a situation in which one processor performs some of the cited functions and other processor(s) perform other parts of the cited functions, and may also cover, but are not limited to, a situation in which a single processor can perform all of the cited functions. Additionally, at least one processor (1970) may include a combination of processors performing the various cited / disclosed functions in a distributed manner, for example. At least one processor (1970) may execute program instructions to achieve or perform the various functions.
[0218] According to one embodiment, at least one processor (1970) may include at least one of a CPU (central processing unit), NPU (graphics processing unit), MPU (micro processing unit), MCU (micro controller unit), AP (application processor), CP (communication processor), SoC (system on chip), or IC (integrated circuit) sensor hub, supplementary processor, ASIC (application specific integrated circuit), or FPGA (field programmable gate arrays), and may have multiple cores.
[0219] According to one embodiment, a memory (1980) (e.g., memory (130) of FIG. 1a) may store various data that can be used to control the operation of each component of the first electronic device (1400). The memory (1980) may include, for example, at least one storage medium that stores a plurality of applications used in the first electronic device (1900), data for controlling the operation of the first electronic device (1900), and instructions. When the instructions stored in the memory (1980) are executed on at least one processor (1970), they may cause the first electronic device (1900) to perform at least one operation (e.g., at least one of the operations described above in FIG. 1 to 18).
[0220] According to one embodiment, the memory (1980) may store at least one program for processing and controlling the processor (1970) and may store input and / or output data. The memory (1980) may also store at least one artificial intelligence model. The memory (1980) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD (secure digital) or XD (extreme digital) memory), RAM (random access memory), SRAM (static random access memory), ROM (read-only memory), EEPROM (electrically erasable programmable read-only memory), PROM (programmable read-only memory), magnetic memory, a magnetic disk, and an optical disk. According to one example, a web storage or cloud server that performs storage functions on the Internet may be operated by the first electronic device (1900).
[0221] FIG. 20 is a drawing illustrating the configuration of a second electronic device according to one embodiment of the present disclosure.
[0222] According to one embodiment, the second electronic device (2000) may be an electronic device connected to or capable of communicating with the first electronic device (e.g., the first electronic device (1900) of FIG. 19). The second electronic device (2000) may be, for example, the wearable electronic device (102) of FIG. 1a or the wearable electronic device (400) of FIG. 4.
[0223] Referring to FIG. 20, according to one embodiment, the second electronic device (2000) may include, but is not limited to, at least one camera (2010), at least one microphone (2020), at least one speaker (2030), at least one sensor (2040), at least one communication circuit (2050), at least one processor (2060) and / or at least one memory (2070). The second electronic device (2000) may omit some of the components described above and may further include other components necessary to perform at least one operation.
[0224] According to one embodiment, at least one processor (2060) (e.g., processor (120) of FIG. 1a) may be electrically and / or operationally connected to other components (e.g., memory (2070)) by a configuration such as a communication bus.
[0225] According to one embodiment, at least one camera (2010) (e.g., camera module (180) of FIG. 1a) can capture still images and video. According to one embodiment, at least one camera (2010) may include one or more lenses, image sensors, image signal processors, or flashes.
[0226] According to one embodiment, at least one microphone (2020) (e.g., input module (150) of FIG. 1a) can receive a sound signal (or audio signal) containing commands or data to be used by a component (e.g., processor (2060)) of the second electronic device (2000) from outside the second electronic device (2000).
[0227] According to one embodiment, at least one speaker (2030) (e.g., the acoustic output module (155) of FIG. 1a) can output an acoustic signal to the outside of the second electronic device (2000).
[0228] According to one embodiment, at least one sensor (2040) (e.g., sensor module (176) of FIG. 1a) can detect the operating state (e.g., power or temperature) of the second electronic device (2000) or an external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, at least one sensor (2040) may include, for example, an inertial sensor, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0229] According to one embodiment, at least one communication circuit (2050) (e.g., communication module (190) of FIG. 1a) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between a second electronic device (2000) and an external electronic device, and the performance of communication through the established communication channel.
[0230] According to one embodiment, at least one processor (2060) can control the overall operation of the second electronic device (2000) and can perform at least one operation of the second electronic device (2000) (e.g., at least one of the operations described above in FIG. 1a through 18). The processor (2060) can perform operations or data processing regarding the control and / or communication of at least one other component of the second electronic device (2000). The processor (2060) may include at least one processing circuit that executes instructions stored in memory (2070).
[0231] According to one embodiment, at least one processor (2060) may include various processing circuits and / or multiple processors. One or more of the at least one processor (2060) may be configured to perform various functions described in the present disclosure individually and / or collectively. Where the “processor,” “at least one processor,” and “one or more processors” are described in the present disclosure as being configured to perform various functions, these terms may cover, for example, a situation in which one processor performs some of the cited functions and other processor(s) perform other parts of the cited functions, and may also cover, but are not limited to, a situation in which a single processor can perform all of the cited functions. Additionally, at least one processor (2060) may include a combination of processors performing the cited / disclosed various functions, for example, in a distributed manner. At least one processor (2060) may execute program instructions to achieve or perform various functions.
[0232] According to one embodiment, at least one processor (2060) may include at least one of a CPU, NPU, GPU, MPU, MCU, AP, CP, SoC, or IC sensor hub, auxiliary processor, ASIC, or FPGA, and may have multiple cores.
[0233] According to one embodiment, a memory (2070) (e.g., memory (130) of FIG. 1a) may store various data that can be used to control the operation of each component of the second electronic device (2000). The memory (2070) may include, for example, at least one storage medium that stores a plurality of applications used in the second electronic device (2000), data for controlling the operation of the second electronic device (2000), and instructions. When the instructions stored in the memory (2070) are executed on at least one processor (2060), they may cause the second electronic device (2000) to perform at least one operation (e.g., at least one of the operations described above in FIG. 1 to 18).
[0234] According to one embodiment, the memory (2070) may store at least one program for processing and controlling the processor (2060) and may store input and / or output data. The memory (2070) may also store at least one artificial intelligence model. The memory (2070) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory), RAM, SRAM, ROM, EEPROM, PROM, magnetic memory, a magnetic disk, or an optical disk. According to one example, a web storage or cloud server that performs storage functions on the Internet may be operated by the second electronic device (2000).
[0235] The present disclosure provides a method for providing a user with meaningful keywords or scene information among ambient sounds without missing important notifications in an audiovisual immersion situation, such as playing a game or performing work, while wearing, for example, an HMD device (e.g., a VST device) and earbuds.
[0236] The present disclosure can help users better understand the context of a notification by, for example, generating and providing video and specific descriptive information about a sound source through a generative AI model based on user meta-information.
[0237] According to one embodiment of the present disclosure, a head-mounted device (HMD) device (101; 200; 300; 1900) may include at least one processor (120; 1970) including a processing circuit; and a memory (130; 1980) including at least one storage medium for storing instructions. When the above commands are executed individually or collectively by the at least one processor, the HMD device: detects a sound source that provides meaningful sound information from a sound signal acquired through one or more microphones based on at least one of location information or behavior information of a user wearing the HMD device; acquires first image information regarding the sound source through a camera located in the direction of the sound source based on direction information of the sound source; identifies whether a sound source object corresponding to the sound source is included within the image of the first image information; generates first descriptive information regarding the sound source using the first image information and user-related information based on the identification that the sound source object is included within the image; generates second image information including the sound source object based on the identification that the sound source object is not included within the image; generates second descriptive information regarding the sound source using the second image information and user-related information; and the first descriptive information or the second descriptive information It may cause to provide notification information that includes
[0238] According to one embodiment, the second image information is generated using a generative AI model, and the generative AI model is configured to receive prompt data generated based on information about the sound source and information related to the user as input data and output the second image information, wherein the information about the sound source includes information about an estimated sound source object, and the information about the estimated sound source object may be generated based on correlation analysis between at least one candidate object within the area where the user is located and the sound signal.
[0239] According to one embodiment, the first description information and the second setting information are generated using a generative AI model, and the generative AI model may be configured to receive prompt data generated based on the first image information and the user-related information as input data and output the first description information, or receive prompt data generated based on the second image information and the user-related information as input data and output the second description information.
[0240] According to one embodiment, the user-related information may include at least one of the user's identification information, the user's schedule information, the user's preference information, or information regarding the usage history of an application used by the user.
[0241] According to one embodiment, the operation of detecting the sound source may include: an operation of updating at least one reference keyword and at least one reference scene token corresponding to the user's situation based on at least one of the location information or the action information.
[0242] According to one embodiment, the operation of detecting the sound source may include: an operation of performing keyword matching between the at least one keyword and the at least one reference keyword in response to at least one keyword being obtained from the sound signal; an operation of performing scene matching between the at least one scene token and the at least one reference scene token in response to at least one scene token being obtained from the sound signal; and an operation of determining whether a sound source providing meaningful sound information is detected based on the result of the keyword matching and the scene matching.
[0243] According to one embodiment, the direction information of the sound source is obtained based on time difference of arrival (TDOA) information indicating the difference in time when the sound signal arrives through each of a plurality of microphones, and the plurality of microphones may be included in at least one of the HMD device or a wearable electronic device (102; 400; 2000) connected to the HMD device.
[0244] According to one embodiment, the operation of acquiring the first image information may include: selecting a first camera located in the direction of the sound source among a plurality of cameras using the direction information of the sound source; activating the selected first camera; and acquiring the first image information of the sound source through the activated first camera. The first camera may be included in the HMD device or a wearable electronic device connected to the HMD device.
[0245] According to one embodiment, the HMD device includes a communication circuit (190; 1960), and when the first camera is included in the wearable electronic device: the operation of activating the selected first camera includes the operation of transmitting a command to activate the first camera to the wearable electronic device through the communication circuit, and the operation of acquiring the first image information may include the operation of receiving the first image information captured through the first camera from the wearable device through the communication circuit.
[0246] According to one embodiment, the first camera may be activated only for a specified period.
[0247] The embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0248] The term “module” as used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0249] One embodiment of the present document may be implemented as software (e.g., program (140) of FIG. 1) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) of FIG. 1 or external memory (138) of FIG. 1) that is readable by a machine (e.g., electronic device (101) of FIG. 1). For example, a processor (e.g., processor (120) of FIG. 1) of the machine (e.g., electronic device (101) of FIG. 1) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0250] According to one embodiment, the method according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0251] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In a head-mounted device (HMD) device (101; 200; 300; 1900), At least one processor (120;1970) including a processing circuit; and The device includes a memory (130; 1980) comprising at least one storage medium for storing instructions, wherein the instructions, when executed individually or collectively by the at least one processor, cause the HMD device: Based on at least one of location information or behavior information of a user wearing the above HMD device, a sound source to which explanatory information is to be provided is detected from a sound signal obtained through one or more microphones, and Based on the direction information of the sound source, first image information of the sound source is obtained through a camera located in the direction of the sound source, and Identify whether a sound source object corresponding to the sound source is included within the image of the first image information, and Based on the identification that the sound source object is included within the above image, first descriptive information for the sound source is generated using the first image information and user-related information, and Based on the identification that the sound source object is not included in the above video, a second video information including the sound source object is generated, and second descriptive information for the sound source is generated using the second video information and the user-related information. An HMD device that causes to provide notification information including the first explanatory information or the second explanatory information.
2. In Paragraph 1, The above second image information is generated using a generative AI (artificial intelligence) model, and The above generative AI model is configured to receive prompt data generated based on information about the sound source and information related to the user as input data, and to output the second video information. An HMD device wherein information regarding the sound source includes information regarding an estimated sound source object, and the information regarding the estimated sound source object is generated based on correlation analysis between at least one candidate object within the area where the user is located and the sound signal.
3. In Paragraph 1 or 2, The above first explanatory information and the above second setting information are generated using a generative AI model, and An HMD device configured such that the generative AI model receives prompt data generated based on the first image information and the user-related information as input data and outputs the first explanation information, or receives prompt data generated based on the second image information and the user-related information as input data and outputs the second explanation information.
4. In Paragraph 2 or 3, An HMD device wherein the user-related information comprises at least one of the user's identification information, the user's schedule information, the user's preference information, or information regarding the usage history of an application used by the user.
5. In any one of paragraphs 1 through 4, The operation of detecting the above sound source is: An HMD device comprising an operation to update at least one reference keyword and at least one reference scene token corresponding to the user's situation based on at least one of the above location information or the above action information.
6. In Paragraph 5, The operation of detecting the above sound source is: An operation to perform keyword matching between the at least one keyword and the at least one reference keyword in response to obtaining at least one keyword from the sound signal; An operation to perform scene matching between the at least one scene token and the at least one reference scene token in response to obtaining at least one scene token from the sound signal; and An HMD device comprising an operation to determine whether a sound source providing meaningful sound information is detected based on the results of the keyword matching and scene matching.
7. In any one of paragraphs 1 through 6, The direction information of the above sound source is, The above sound signal is acquired based on time difference of arrival (TDOA) information indicating the time difference when each of the multiple microphones arrives, and The above plurality of microphones are, An HMD device included in at least one of the above HMD device or a wearable electronic device (102; 400; 2000) connected to the above HMD device.
8. In Paragraph 7, The operation of acquiring the above-mentioned first image information is: An operation of selecting a first camera located in the direction of the sound source among a plurality of cameras using the direction information of the sound source; The operation of activating the first camera selected above; and The method includes the operation of acquiring first image information about the sound source through the activated first camera, and The first camera is an HMD device included in the HMD device or a wearable electronic device connected to the HMD device.
9. In Paragraph 8, It includes a communication circuit (190;1960), When the first camera is included in the wearable electronic device: The operation of activating the selected first camera includes the operation of transmitting a command to activate the first camera to the wearable electronic device through the communication circuit, and An HMD device, wherein the operation of acquiring the first image information includes the operation of receiving the first image information captured through the first camera from the wearable device through the communication circuit.
10. In Paragraph 8 or 9, The above-mentioned first camera is an HMD device that is activated only for a specified period.
11. In a method for an HMD (head mounted display) device, An operation of detecting a sound source to which explanatory information is to be provided from a sound signal obtained through one or more microphones, based on at least one of location information or behavior information of a user wearing the above HMD device; An operation of acquiring first image information of the sound source through a camera positioned in the direction of the sound source, based on the direction information of the sound source; An operation to identify whether a sound source object corresponding to the sound source is included within the image of the first image information; An operation to generate first descriptive information about the sound source using the first image information and user-related information, based on the identification that the sound source object is included in the above image; Based on identifying that the sound source object is not included in the above image, generating second image information including the sound source object, and generating second description information for the sound source using the second image information and the user-related information; and A method comprising an operation of providing notification information including the first explanatory information or the second explanatory information.
12. In Paragraph 11, The above second image information is generated using a generative AI model, and The above generative AI model is configured to receive prompt data generated based on information about the sound source and user-related information as input data, and to output the second video information. A method in which information regarding the sound source includes information regarding an estimated sound source object, and is generated based on a correlation analysis of at least one candidate object within the area where the user is located and the sound signal.
13. In Paragraph 11 or 12, The above first explanatory information and the above second setting information are generated using a generative AI model, and A method in which the generative AI model is configured to receive prompt data generated based on the first image information and user-related information as input data and output the first explanation information, or receive prompt data generated based on the second image information and user-related information as input data and output the second explanation information.
14. In Paragraph 12 or 13, A method comprising at least one of the user-related information, the user's identification information, the user's schedule information, the user's preference information, or information regarding the usage history of an application used by the user.
15. In any one of paragraphs 11 through 14, The operation of detecting the above sound source is: A method comprising the operation of updating at least one reference keyword and at least one reference scene token corresponding to the user's situation based on at least one of the user's location information or behavior information.