Electronic device, method, and computer-readable storage medium for using multimodal model

Multimodal models within electronic devices enhance voice recognition by integrating audio and visual data to improve accuracy in noisy and low-volume environments, addressing existing challenges in voice command recognition.

WO2025249802A1PCT designated stage Publication Date: 2025-12-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006520
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-08
Filing Date
2025-05-14
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in accurately recognizing voice commands in noisy environments and low-volume speech scenarios, leading to suboptimal voice recognition performance.

Method used

The implementation of multimodal models within electronic devices that utilize both audio and image data, including sensors and microphones, to enhance voice recognition by identifying reference gestures and voice commands, and adapt to environmental noise levels.

Benefits of technology

Improves voice recognition accuracy in noisy and low-volume environments by leveraging multimodal models to integrate visual cues and environmental context, enhancing user interaction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006520_04122025_PF_FP_ABST
    Figure KR2025006520_04122025_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: an image sensor; a microphone; a communication circuit; a memory for storing instructions; and at least one processor including a processing circuit. The instructions, when executed by the at least one processor, may cause the electronic device to: detect an event; acquire a first image; acquire a first audio signal; transmit, to an external electronic device, a signal for waking up a second multimodal model on the basis of identifying, using a first multimodal model, the first audio signal including a reference voice command and the first image representing a reference gesture; after transmitting the signal, acquire a second audio signal and acquire a second image; transmit, to the external electronic device, data acquired using the second audio signal and data about a visual object in the second image corresponding to the lips of a user; receive response information of the second audio signal; and perform a function.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and computer-readable storage medium for utilizing a multimodal model

[0001] The following descriptions relate to electronic devices, methods, and computer-readable storage media for utilizing multimodal models.

[0002] An electronic device may include a microphone. The electronic device may acquire an audio signal through the microphone. For example, the audio signal may include speech or voice uttered by a speaker. The electronic device may provide a function for recognizing the voice within the audio signal.

[0003] The electronic device may include a communication circuit. The electronic device may transmit data to and / or receive data from an external electronic device via the communication circuit.

[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0005] An electronic device is provided. The electronic device may include an image sensor facing a front side of the electronic device. The electronic device may include a microphone. The electronic device may include communication circuitry. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for driving the image sensor. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first image through the image sensor driven in response to the event, and to acquire a first audio signal through the microphone, based on the detection. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, a signal to an external electronic device, wherein the signal causes a wake-up of a second multimodal model within the external electronic device based on identifying, using a first multimodal model within the electronic device, the first image representing a reference gesture of the user and the first audio signal including a reference voice command. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device, after transmitting the signal, to acquire a second audio signal via the microphone and to acquire a second image via the image sensor.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, to the external electronic device via the communication circuit, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device via the communication circuit, response information regarding the second audio signal, obtained using the second multimodal model while in a wake-up state in response to the signal. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.

[0006] A method is provided. The method can be executed within an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit. The method can include detecting an event for driving the image sensor. The method can include acquiring a first image through the image sensor driven according to the event and acquiring a first audio signal through the microphone, based on the detection. The method can include transmitting a signal to an external electronic device via the communication circuit, the signal causing a wake-up of a second multimodal model within the external electronic device, based on identifying the first image expressing a reference gesture of a user and the first audio signal including a reference voice command, using a first multimodal model within the electronic device. The method can include, after transmitting the signal, acquiring a second audio signal through the microphone and acquiring a second image through the image sensor. The method may include an operation of transmitting, to the external electronic device, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips, via the communication circuit. The method may include an operation of receiving, from the external electronic device, response information regarding the second audio signal obtained using the second multimodal model in a wake-up state according to the signal, via the communication circuit. The method may include an operation of performing a function according to the response information.

[0007] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit, cause the electronic device to detect an event to drive the image sensor. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image through the image sensor driven according to the event, and to acquire a first audio signal through the microphone, based on the detection. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuitry, a signal to an external electronic device, the signal causing a wake-up of a second multimodal model within the external electronic device based on identifying, using a first multimodal model within the electronic device, the first image representing a reference gesture of the user and the first audio signal including a reference voice command. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, after transmitting the signal, acquire a second audio signal through the microphone and acquire a second image through the image sensor. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuitry, first data acquired using the second audio signal and second data regarding a visual object within the second image corresponding to the user's lips, to the external electronic device.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, response information for the second audio signal obtained using the second multimodal model in a wake-up state in response to the signal from the external electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function in response to the response information.

[0008] An electronic device is provided. The electronic device may include an image sensor facing a front side of the electronic device. The electronic device may include a microphone. The electronic device may include communication circuitry. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first audio signal via the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, a signal to the external electronic device, in accordance with a reference voice command included in the first audio signal, to cause a wake-up of a first multimodal model within the external electronic device, based on identifying the type of the environment as a first type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor based on identifying the type of the environment as a second type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first image via the driven image sensor.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a second audio signal via the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, to the external electronic device, the signal causing a wake-up of the first multimodal model based on identifying, using a second multimodal model within the electronic device, the first image representing a reference gesture of the user and the second audio signal including the reference voice command.

[0009] A method is provided. The method may be executed within an electronic device having an image sensor facing a front side of the electronic device, a microphone, and communication circuitry. The method may include acquiring a first audio signal via the microphone. The method may include identifying a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device. The method may include transmitting, via the communication circuitry, a signal to an external electronic device, based on identifying that the type of environment is the first type, a wake-up signal for a first multimodal model within the external electronic device, in accordance with a reference voice command included in the first audio signal. The method may include driving the image sensor based on identifying that the type of environment is the second type. The method may include acquiring a first image via the driven image sensor. The method may include acquiring a second audio signal via the microphone. The method may include an operation of transmitting, to the external electronic device through the communication circuit, a signal causing a wake-up of the first multimodal model based on identifying the first image representing the user's reference gesture and the second audio signal including the reference voice command using a second multimodal model within the electronic device.

[0010] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and communication circuitry, cause the electronic device to acquire a first audio signal via the microphone. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, from the first audio signal, a type of environment in which the electronic device is located, using a model within the electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, via the communication circuitry, a signal to an external electronic device, the signal causing a wake-up of a first multimodal model within the external electronic device, based on identifying that the type of environment is the first type, in accordance with a reference voice command included in the first audio signal. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to drive the image sensor based on identifying that the type of the environment is a second type. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image via the driven image sensor. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a second audio signal via the microphone.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuit, to the external electronic device a signal that causes a wake-up of the first multimodal model based on identifying, using a second multimodal model within the electronic device, the first image representing the user's reference gesture and the second audio signal including the reference voice command.

[0011] An electronic device is provided. The electronic device may include an image sensor facing a front side of the electronic device. The electronic device may include a microphone. The electronic device may include communication circuitry. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for a call connection with an external electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first audio signal through the microphone based on performing the call connection with the external electronic device based on the detection, and to identify a type of environment in which the electronic device is located from the first audio signal through the microphone using a model within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a second audio signal through the microphone and transmit the second audio signal to the external electronic device through the communication circuitry based on identifying that the type of the environment is a first type.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor, acquire an image through the driven image sensor, acquire a third audio signal through the microphone, and provide first data acquired using the third audio signal and second data about a visual object in the image corresponding to the user's lips to a multimodal model within the electronic device, thereby obtaining response information about the third audio signal and the visual object, and performing a function according to the response information.

[0012] A method is provided. The method can be executed within an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit. The method can include detecting an event for a call connection with an external electronic device. The method can include, based on the detection, performing the call connection with the external electronic device, acquiring a first audio signal through the microphone, and using a model within the electronic device, identifying a type of environment in which the electronic device is located from the first audio signal through the microphone. The method can include, based on identifying that the type of the environment is the first type, acquiring a second audio signal through the microphone, and transmitting the second audio signal to the external electronic device through the communication circuit. The method may include an operation of driving the image sensor based on identifying that the type of the environment is the second type, acquiring an image through the driven image sensor, acquiring a third audio signal through the microphone, and providing first data acquired using the third audio signal and second data about a visual object in the image corresponding to the user's lips to a multimodal model in the electronic device, thereby acquiring response information about the third audio signal and the visual object, and performing a function according to the response information.

[0013] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device having an image sensor, a microphone, and a communication circuit facing a front side of the electronic device, cause the electronic device to detect an event for a call connection with an external electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first audio signal through the microphone based on the detection of performing the call connection with the external electronic device, and to identify a type of environment in which the electronic device is located from the first audio signal through the microphone using a model within the electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a second audio signal via the microphone and transmit the second audio signal to the external electronic device via the communication circuitry based on identifying that the type of the environment is a first type.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to drive the image sensor, acquire an image through the driven image sensor, acquire a third audio signal through the microphone, and provide first data acquired using the third audio signal and second data about a visual object in the image corresponding to the user's lips to a multimodal model within the electronic device, thereby obtaining response information about the third audio signal and the visual object, and performing a function according to the response information.

[0014] A wearable device is provided. The wearable device may include at least one display. The wearable device may include a microphone. The wearable device may include at least one sensor. The wearable device may include a memory including one or more storage media storing instructions. The wearable device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to display, through the at least one display, a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through pass-through. The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen. The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to identify, through the at least one sensor, a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal. The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to identify, based on the first direction of the gaze of the user and the second direction of the gaze of the visual object, whether the user and a speaker corresponding to the visual object are looking at each other.The instructions, when individually or collectively executed by the at least one processor, may cause the wearable device to display, through the at least one display, text generated from the voice signal, superimposed on the screen, based on identifying that the user and the other user corresponding to the visual object are viewing each other.

[0015] A method is provided. The method can be performed in a wearable device having at least one display, a microphone, and at least one sensor. The method can include an operation of displaying a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through a pass-through through the at least one display. The method can include an operation of identifying a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen based on an audio signal acquired through the microphone while the screen is displayed. The method can include an operation of identifying a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal through the at least one sensor. The method can include an operation of identifying whether the user and a speaker corresponding to the visual object are looking at each other based on the first direction of the gaze of the user and the second direction of the gaze of the visual object. The method may include an action of displaying text generated from the voice signal by overlaying it on the screen through the at least one display based on identifying that the user and the other user corresponding to the visual object are seeing each other.

[0016] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by a wearable device having at least one display, a microphone, and at least one sensor, cause the wearable device to display, through the at least one display, a screen representing at least a portion of a three-dimensional space corresponding to a physical environment through pass-through. The one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, based on an audio signal acquired through the microphone while displaying the screen, a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in the screen. The one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, through the at least one sensor, a first direction of a gaze of a user wearing the wearable device and a second direction of a gaze of the visual object corresponding to the voice signal. The one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to identify, based on the first direction of the gaze of the user and the second direction of the gaze of the visual object, whether the user and a speaker corresponding to the visual object are looking at each other.The one or more programs may include instructions that, when executed by the wearable device, cause the wearable device to display, through the at least one display, text generated from the voice signal, superimposed on the screen, based on identifying that the user and the other user corresponding to the visual object are viewing each other.

[0017] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0018] Figure 1 illustrates an example of an environment including an electronic device and an external electronic device.

[0019] Figure 2a illustrates an example of an environment where the volume of noise including electronic devices is greater than the user's voice volume.

[0020] Figure 2b illustrates an example of an environment in which a user including an electronic device must whisper utterances.

[0021] Figure 3a is a simplified block diagram of an exemplary electronic device.

[0022] FIG. 3b is a simplified block diagram of another exemplary electronic device and external electronic device.

[0023] Figure 4a shows an example of learned actions of the first multimodal model.

[0024] Figure 4b shows an example of learned actions of the fourth multimodal model.

[0025] Figure 5 illustrates examples of operations in which an electronic device transmits a signal that causes a wake-up of a second multimodal model.

[0026] Figures 6a and 6b illustrate examples of events for driving at least one sensor.

[0027] Figure 7 illustrates examples of reference gestures and reference voice commands recognized by the first multimodal model.

[0028] Figures 8a and 8b illustrate examples of operations in which an electronic device transmits a signal to cause a wake-up of a second multimodal model.

[0029] Figures 9a and 9b illustrate examples of operations in which an electronic device transmits data to an external electronic device and receives response information from the external electronic device to utilize a second multimodal model.

[0030] Figures 10a and 10b illustrate examples of electronic devices that perform functions according to response information.

[0031] Figure 11 illustrates an example in which an electronic device performs a call connection with an external electronic device.

[0032] FIG. 12 is a block diagram of an electronic device within a network environment according to various embodiments.

[0033] Figure 13a illustrates an example of a perspective view of a wearable device.

[0034] FIG. 13b illustrates an example of one or more hardware devices arranged within a wearable device.

[0035] FIG. 13c illustrates an example of a wearable device according to one embodiment.

[0036] Figures 14a and 14b illustrate an example of the appearance of a wearable device.

[0037] Figure 15 illustrates an example of the appearance of a wearable device.

[0038] Figure 16 illustrates an example of a block diagram of a wearable device.

[0039] Figure 17 illustrates an example block diagram of a wearable device for displaying images in virtual space.

[0040] Figure 18 illustrates an example of a wearable device that recognizes audio signals emitted from external objects.

[0041] Figure 19 illustrates an example of a wearable device that adaptively provides a response to an audio signal based on user input.

[0042] FIGS. 20A and 20B illustrate examples of wearable devices that adaptively provide responses to audio signals based on a user's gaze.

[0043] Figure 21 illustrates an example of a wearable device that provides a response in a second language to an audio signal spoken in a first language.

[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0045] The terms used in this disclosure are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.

[0046] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0047] In the following description, terms referring to data (e.g., data, information, sensing data, signal), terms referring to values ​​(e.g., threshold value, reference value), terms for operational states (e.g., operation, process), terms referring to objects (e.g., visual objects, emoji graphical objects), terms referring to network entities, terms referring to components of devices, etc. are examples for convenience of explanation. Therefore, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used. In addition, terms such as '... part', '... device', '... object', '... body', etc. used below may mean at least one shape structure or a unit that processes a function.

[0048] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." A condition described as "more than" may be replaced with "more than," a condition described as "less than" may be replaced with "less than," and a condition described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of elements from A (including A) to B (including B). hereinafter, "C" and / or "D" mean at least one of "C" or "D," that is, including {"C", "D", "C" and "D"}.

[0049] Figure 1 illustrates an example of an environment including an electronic device and an external electronic device.

[0050] Referring to FIG. 1, the electronic device (120) may include a microphone. For example, the microphone may operate in a state for low power consumption. For example, the microphone may detect an audio signal (or sound signal) while in the state for low power consumption. For example, the electronic device (120) may acquire an audio signal (or sound signal) through the microphone while the microphone is in the state for low power consumption. For example, the electronic device (120) may acquire an audio signal from an external environment through the microphone. For example, the audio signal may include a voice spoken by a user (110). For example, the audio signal may include noise.

[0051] For example, the electronic device (120) can recognize the audio signal. For example, the electronic device (120) can recognize a voice signal within the audio signal. For example, the electronic device (120) can provide a voice recognition function. For example, the electronic device (120) can obtain data using the recognized audio signal. For example, the data can correspond to the audio signal. For example, the data can include an audio signal that has undergone post-processing on the audio signal. For example, the post-processing can be performed to input the audio signal to a model (e.g., an artificial intelligence model). For example, the data can include text representing the audio signal. For example, the data can include text representing a voice signal (or voice signal) included in the audio signal.

[0052] For example, the electronic device (120) may use a model for a voice recognition function. For example, the electronic device (120) may train (or learn) the model. For example, when the electronic device (120) recognizes a voice signal using the trained model, the voice signal may be recognized more accurately than when the electronic device (120) recognizes the voice signal without using the trained model. For example, the electronic device (120) may be connected to another model in an external electronic device (130) to use the voice recognition function. For example, the electronic device (120) may include a communication circuit. For example, the electronic device (120) may transmit the audio signal to the external electronic device (130) through the communication circuit. For example, the electronic device (120) may transmit data acquired using the audio signal to the external electronic device (130) through the communication circuit.

[0053] For example, the electronic device (120) may include a smartphone. For example, the electronic device (120) may include a wearable device. For example, the electronic device (120) may include a smartwatch. As a non-limiting example, the electronic device (120) may include a portable electronic device such as a smartphone, a tablet, a laptop computer, or a smartwatch. For example, the electronic device (120) may be described as a multi-function device or a user device.

[0054] For example, the external electronic device (130) can recognize a voice signal. For example, the external electronic device (130) can use a model for recognizing the voice signal. For example, the external electronic device (130) can train the model. For example, the external electronic device (130) can recognize the voice signal more accurately by using the trained model. For example, the model of the external electronic device (130) can be more complex than the model of the electronic device (120).

[0055] For example, the external electronic device (130) can output (or obtain) response information for the recognized voice signal. For example, the external electronic device (130) can recognize the voice signal using the model, generate a prompt based on the voice signal, and provide the prompt to a large language model within the external electronic device (130). For example, the external electronic device (130) can output or obtain the response information for the voice signal using the large language model.

[0056] For example, the external electronic device (130) may include a communication circuit. For example, the external electronic device (130) may receive a signal (or data) from the electronic device (120) through the communication circuit. For example, the external electronic device (130) may receive a signal from the electronic device (120) that causes a wake-up of a model within the external electronic device (130). For example, the signal may change the model in a state for low power consumption to an active state. For example, the signal may be referred to as a signal that operates the model in the state for low power consumption.

[0057] For example, the external electronic device (130) can receive or obtain a voice signal from the electronic device (120). For example, the external electronic device (130) can receive or obtain data obtained using the voice signal from the electronic device (120). For example, the external electronic device (130) can recognize the obtained voice signal using the model and generate a prompt. For example, the external electronic device (130) can obtain response information by providing the prompt to the large language model.

[0058] For example, within state (100), a user (110) may utter a reference voice command (or wake word) to activate a voice recognition function of an electronic device (120). For example, the electronic device (120) may acquire an audio signal including the reference voice command via the microphone. For example, the electronic device (120) may cause the operation of the voice recognition function based on noise within the acquired audio signal and the reference voice command.

[0059] Figure 2a illustrates an example of an environment where the volume of noise including electronic devices is greater than the user's voice volume.

[0060] Referring to FIG. 2a, the environment (200) can be described as an environment with a relatively high noise volume. For example, the environment (200) may include an indoor environment. For example, the environment (200) may include an outdoor environment. As non-limiting examples, the environment (200) may include a cafe, a restaurant, a subway, a construction site, a plaza, or a concert hall, but the embodiments are not limited thereto.

[0061] For example, the environment (200) may represent an environment in which the signal-to-noise ratio (SNR) of an audio signal acquired through the microphone of the electronic device (120) is below a reference value. For example, the sound quality of the audio signal in the environment (200) may be relatively low. For example, the environment (200) may represent an environment in which the volume of the user's (110) voice is low in relation to the noise volume. For example, the higher the volume of noise in the environment (200), the lower the quality of the voice recognition function of the electronic device (120).

[0062] Figure 2b illustrates an example of an environment in which a user including an electronic device must whisper utterances.

[0063] Referring to FIG. 2b, the environment (210) may represent an environment in which the user (110) cannot speak at a relatively high volume. For example, the environment (210) may represent a library. For example, the environment (210) is not limited to a library. The environment (210) may include, but is not limited to, a conference room, a movie theater, or an art gallery.

[0064] For example, in an environment (210), a user (110) may whisper to an electronic device (120). For example, in an environment (210), the voice volume of the user (110) may be lower than a reference volume that the electronic device (120) can recognize. For example, because the voice volume of the user (110) is lower than the reference volume, the recognition quality of the electronic device (120) may deteriorate. For example, in an environment (210), the electronic device (120) may not recognize the voice of the user (110).

[0065] For example, since the quality of the voice recognition function is relatively poor in the environment (200) of FIG. 2A, a method for removing noise from an audio signal received through a microphone of an electronic device (120) and / or a method for enhancing a voice signal within the audio signal may be required. For example, a method for assisting voice recognition using at least one sensor of the electronic device (120) may be required.

[0066] For example, in the environment (210) of FIG. 2B, since the voice volume of the user (110) is lower than the reference volume that the electronic device (120) can recognize, a method for recognizing a reference gesture may be required to operate the voice recognition function of the electronic device (120). For example, in the environment (210), a method for recognizing the shape of the lips of the user (110) may be required. For example, in the environment (210), a method for identifying whether sensing data acquired through at least one sensor of the electronic device (120) corresponds to the reference sensing data may be required.

[0067] These methods may be implemented within an electronic device, as exemplified below. For example, the electronic device exemplified below may include components (or hardware components) for providing these methods. These components are described and exemplified in more detail with reference to FIG. 3A.

[0068] FIG. 3A is a simplified block diagram of an exemplary electronic device. The electronic device (301) may be an example of the electronic device (101) of FIG. 1.

[0069] Referring to FIG. 3A, the electronic device (301) may include at least one processor (300), a communication circuit (310), a memory (320), a microphone (330), at least one sensor (340), a display (311), and / or a speaker (312). For example, the at least one processor (300), the communication circuit (310), the memory (320), the microphone (330), at least one sensor (340), the display (311), and / or the speaker (312) may be electronically and / or operably coupled with each other by a communication bus. Hereinafter, operably coupled hardware components may mean that a direct connection or an indirect connection is established between the hardware components, either wired or wireless, such that a second hardware component is controlled by a first hardware component among the hardware components. Although the hardware components illustrated in FIG. 3A are illustrated based on different blocks, the present disclosure is not limited thereto. For example, some of the hardware components illustrated in FIG. 3A (e.g., at least one processor (300), a communication circuit (310), and at least a portion of the memory (320)) may be included in a single integrated circuit such as a system on chip (SoC) or a system in package (SIP). The type and / or number of hardware components included in the electronic device (301) are not limited to those illustrated in FIG. 3A. For example, the electronic device (301) may include only some of the hardware components illustrated in FIG. 3A.

[0070] At least one processor (300) may include a hardware component for processing data based on executing instructions. The hardware component for processing data may include, for example, a central processing unit (CPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a graphic processing unit (GPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a display processing unit (DPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a neural processing unit (NPU) (e.g., including processing circuitry).

[0071] At least one processor (300) may include one or more cores. For example, at least one processor (300) may have a multi-core processor architecture, such as a dual core, quad core, or hexa core. The content of the processor (1220) of FIG. 12 may be substantially identically applied to at least one processor (300).

[0072] The communication circuit (310) may include hardware components for supporting transmission and / or reception of signals between the electronic device (301) and the external electronic device (302). The communication circuit (310) may include, for example, at least one of a modem (modulator and demodulator), an antenna, and an optical / electronic (O / E) converter. The communication circuit (310) may support transmission and / or reception of electrical signals based on various types of protocols, such as Ethernet, a local area network (LAN), a wide area network (WAN), wireless fidelity (WiFi), Bluetooth, Bluetooth low energy (BLE), zigbee, long term evolution (LTE), and 5G new radio (NR). Specific details regarding the communication circuit (310) of FIG. 3A may be substantially identically applied to the communication module (1290) and / or the antenna module (1297) of FIG. 12.

[0073] The memory (320) may include a hardware component for storing data and / or instructions input to and / or output from at least one processor (300). For example, the instructions may represent operations and / or actions to be performed on data by at least one processor (300) of the electronic device (301). The memory (320) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, at least one of a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a hard disk, a compact disk, and an embedded multimedia card (EMMC). The specific details of the memory (320) of FIG. 3A may be substantially identically applied to the memory (1230) of FIG. 12.

[0074] The microphone (330) may be configured to acquire audio signals generated around the electronic device (301). For example, the microphone (330) may be used to acquire voice signals within the audio signals. The specific details of the microphone (330) of FIG. 3A may be substantially identical to those of the input module (1250) of FIG. 12.

[0075] At least one sensor (340) may include a hardware component of an electronic device (301) used to acquire an external signal. For example, at least one sensor (340) may detect an external signal and generate an electrical signal or data value corresponding to a detected state. For example, at least one sensor (340) may include an image sensor. For example, at least one sensor (340) may include a heart rate sensor. For example, at least one sensor (340) may include an acceleration sensor. For example, at least one sensor (340) may include a gyro sensor. However, the present invention is not limited thereto. The specific details of at least one sensor (340) of FIG. 3A may be substantially identically applied to the details of the sensor module (1276) of FIG. 12.

[0076] For example, the electronic device (301) may acquire sensing data (e.g., an image (e.g., acquired via an image sensor), heart rate data (e.g., acquired via a heart rate sensor), and / or movement data (e.g., acquired via an acceleration sensor and / or a gyro sensor)) via at least one sensor (340).

[0077] The display (311) may include hardware components of the electronic device (301) used to display a screen. For example, the display (311) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light emitting diode (OLED) or a micro LED. However, the present invention is not limited thereto. For example, the display (311) may include a liquid crystal display (LCD). The specific details of the display (311) of FIG. 3A may be substantially identically applied to the details of the display module (1260) of FIG. 12.

[0078] The speaker (312) can be used to output audio signals to the outside of the electronic device (301). For example, the speaker (312) can be used for general purposes, such as multimedia playback or recording playback. The specific details of the speaker (312) of FIG. 3A can be substantially identically applied to the details of the audio output module (1255) of FIG. 12.

[0079] For example, at least one processor (300) may execute operations within the electronic device (301) to wake up a multimodal model within the external electronic device (302) using sensing data acquired through at least one sensor (340) and / or an audio signal acquired through a microphone (330). For example, at least one processor (300) may execute operations within the electronic device (301) to identify a type of environment in which the electronic device (301) is located. For example, at least one processor (300) may execute operations within the electronic device (301) to perform noise canceling on another audio signal acquired through the microphone (330) according to the type of the environment. The operations will be exemplified within the descriptions of FIGS. 5 to 10B.

[0080] FIG. 3b is a simplified block diagram of another exemplary electronic device and external electronic device.

[0081] Referring to FIG. 3b, the electronic device (301) may include a first multimodal model (350), a third model (370), and / or a fourth multimodal model (380).

[0082] The first multimodal model (350) can be used to identify a voice signal included in an audio signal acquired through the microphone (330). For example, the first multimodal model (350) can be used to identify whether a reference voice command is included in the audio signal acquired through the microphone (330). The first multimodal model (350) can identify whether sensing data acquired through at least one sensor (340) (e.g., an image (e.g., acquired through an image sensor), heart rate data (e.g., acquired through a heart rate sensor), and / or movement data (e.g., acquired through an acceleration sensor and / or a gyro sensor)) represents (or includes) reference sensing data. For example, the reference sensing data can include a reference gesture of a user. For example, the first multimodal model (350) can identify whether an image acquired through an image sensor represents a reference gesture. For example, the first multimodal model (350) can identify whether information acquired through a heart rate sensor represents reference heart rate information. For example, the first multimodal model (350) can identify whether information acquired through an acceleration sensor and / or a gyro sensor represents reference movement.

[0083] For example, the first multimodal model (350) can be trained (or learned) using one or more modalities. For example, a modality can be referred to as a type of data. For example, the type of data can include, but is not limited to, text, images, and / or audio signals.

[0084] For example, the first multimodal model (350) may provide an automatic speech recognition (ASR) function. For example, the first multimodal model (350) may be trained to recognize or identify an audio signal acquired through a microphone (330). For example, the first multimodal model (350) may be trained to perform natural language processing on an audio signal acquired through the microphone (330). For example, the first multimodal model (350) may be trained to recognize or identify sensing data acquired through at least one sensor (340). For example, the first multimodal model (350) may be used to acquire data representing a voice signal (or speech signal) included in an audio signal based on an audio signal and / or sensing data. For example, the data representing the voice signal may include text representing the voice signal. For example, the data representing the voice signal may include an audio signal representing the voice signal. Learning of the first multimodal model (350) will be described and illustrated in more detail with reference to FIG. 4a.

[0085] Although FIG. 3B illustrates that the external electronic device (302) provides the output signal of the first multimodal model (350) transmitted from the electronic device (301) to the second multimodal model (360) within the external electronic device (302), this is merely exemplary. For example, the external electronic device (302) may provide the output signal of the first multimodal model (350) transmitted from the electronic device (301) to the large language model (390) within the external electronic device (302). For example, the output signal of the first multimodal model (350) may include response information for an audio signal and / or sensing data.

[0086] As a non-limiting example, the electronic device (301) may provide a response to the audio signal and / or sensing data by executing a function according to the response information using the response information for the audio signal and / or sensing data. For example, the response information for the audio signal and / or sensing data may be used to allow the electronic device (301) to provide a response based on the audio signal and / or sensing data. For example, the response information for the audio signal and / or sensing data may be used to display text representing a voice signal included in the audio signal through a display (e.g., the display (311)). For example, the response information for the audio signal and / or sensing data may be used to output an audio signal representing a voice signal included in the audio signal through a speaker (e.g., the speaker (312)). For example, the response information for the audio signal and / or sensing data may be used to execute a function corresponding to (or linked to) a reference voice command included in the audio signal. For example, response information to audio signals and / or sensing data may be used to execute a function corresponding to (or linked to) a reference gesture represented by the sensing data.

[0087] For example, at least one processor (300) may provide an audio signal acquired through a microphone (330) and sensing data acquired through at least one sensor (340) (e.g., an image (e.g., acquired through an image sensor), heart rate data (e.g., acquired through a heart rate sensor), and / or movement data (e.g., acquired through an acceleration sensor and / or a gyro sensor)) to the first multimodal model (350). For example, at least one processor (300) may obtain response information for the audio signal and the sensing data by using an ASR function of the first multimodal model (350).

[0088] For example, at least one processor (300) may provide an audio signal acquired through a microphone (330) and an image acquired through at least one sensor (340) to a first multimodal model (350). For example, the acquired image may include a visual object corresponding to the lips of a user of the electronic device (301). For example, at least one processor (300) may use an ASR function of the first multimodal model (350) to acquire response information for the audio signal and the visual object in the image corresponding to the lips of the user.

[0089] For example, at least one processor (300) may transmit the acquired response information to an external electronic device (302) via a communication circuit (310). For example, the response information may include data for input to a large language model (390) within the external electronic device (302). For example, the data for input to the large language model (390) may be displayed or used as a prompt.

[0090] For example, at least one processor (300) may execute a function for acquired response information within the electronic device (301). For example, the function for acquired response information may include an operation for displaying text corresponding to the acquired response information on the display (311). For example, the function for acquired response information may include an operation for displaying an emoji corresponding to the acquired response information on the display (311). However, the present invention is not limited thereto.

[0091] The third model (370) can be used to identify (or classify) the type of environment in which the electronic device (301) is located, from the audio signal acquired through the microphone (330). For example, the third model (370) can identify the type of the environment as the first type if the signal-to-noise ratio (SNR) of the audio signal is equal to or greater than a reference value. For example, the third model (370) can identify the type of the environment as the second type if the SNR of the audio signal is less than the reference value.

[0092] For example, the third model (370) can identify the type of the environment. For example, the third model (370) can provide data about the type of the identified environment to the first multimodal model (350). For example, based on identifying the type of the environment as the second type, the third model (370) can provide data about the second type to the fourth multimodal model (380). For example, the third model (370) can be trained using a conformer algorithm, but is not limited thereto.

[0093] For example, the type of environment in which the electronic device (301) is located may include, but is not limited to, an indoor environment, an outdoor environment, the first type of environment, and / or the second type of environment.

[0094] The fourth multimodal model (380) can be used to perform noise cancellation on an audio signal acquired through the microphone (330). For example, the fourth multimodal model (380) can generate an audio signal from which noise is removed from the input audio signal. For example, the fourth multimodal model (380) can perform noise cancellation on the acquired audio signal using the audio signal acquired through the microphone (330), sensing data acquired through at least one sensor (340), and / or data on the type of the environment provided from the third model.

[0095] For example, the fourth multimodal model (380) may be trained to perform noise cancellation on the acquired audio signal. For example, the fourth multimodal model (380) may represent a trained artificial intelligence model. The training of the fourth multimodal model (380) will be described and illustrated in more detail with reference to FIG. 4B.

[0096] The external electronic device (302) may include a second multimodal model (360) and / or a large language model (390).

[0097] The second multimodal model (360) can provide the same functionality as the first multimodal model (350). For example, the second multimodal model (360) can provide an ASR function. For example, the second multimodal model (360) can correspond to the first multimodal model (350). For example, the second multimodal model (360) can be trained (or learned) using one or more modalities. For example, the second multimodal model (360) can be learned using the same algorithm as the first multimodal model (350). For example, the second multimodal model (360) can be more complex than the first multimodal model (350). For example, the second multimodal model (360) can have more layers than the first multimodal model (350). For example, the second multimodal model (360) may have a larger number of parameters than the first multimodal model (350).

[0098] A large language model (390) may be referred to as a language model composed of an artificial neural network that has been pre-trained with a large amount of text data. The large language model (390) may include parameters that are more than 10 times as many as existing general language models (for example, more than 100 billion parameters). The large language model (390) may use a transformer artificial neural network structure based on an attention mechanism. The attention mechanism is a technology that helps an artificial intelligence model to focus on important parts within input data. The attention mechanism may be used to predict output data by predicting the degree to which at least a portion of time-series input data (for example, input data such as voice or video, or input data of some layers of a neural network) contributes to the intermediate or final output of the neural network. The recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, has poor prediction performance when there is information dependency between long time series distances, but the attention mechanism can consider information dependency between long time series distances by controlling the degree of weight concentration within the overall (or partial) context of the input data.

[0099] For example, a large language model (390) may include a transformer with an encoder-decoder structure. The encoder may process input data to output compressed information (e.g., an attention mechanism), and the decoder may process the compressed information to output output data in token units. Each of the encoder and decoder may include an independent attention network, and may include a cross-attention network connecting the encoder and decoder.

[0100] For example, a large language model (390) can be trained in two stages: pre-training and fine-tuning. Pre-training is the process of allowing a large language model (390) to process a large amount of text data and acquire general linguistic knowledge. For example, it can include self-supervised learning to predict the next word using a previous word sequence in a text sequence. Fine-tuning is the process of training a large language model (390) to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. Based on a pre-trained model, additional supervised learning (or adaptive learning) can be performed using a dataset suitable for the domain purpose. A large language model (390) can perform a task with a text input containing natural language called a prompt. For example, a large language model (390) can include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer). The term "LLM (Large Language Model)" can refer to the neural network model itself, but can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot such as chatGPT can also be referred to as an LLM. "LLM" can also include an inference engine that utilizes the LLM neural network model. For example, "entering an input prompt into an LLM" can be referred to as "entering an input prompt into an LLM-based inference engine."

[0101] The large language model (390) is an artificial intelligence model trained with text data and can be used to provide response information for a voice signal within an audio signal. For example, the audio signal may include an audio signal received from an electronic device (301). For example, the voice signal within the audio signal may include a voice signal acquired by a microphone (330) of the electronic device (301). The large language model (390) may be referred to as a large language model and a large language model (LLM). For example, the large language model (390) may perform natural language processing. For example, the large language model (390) may perform natural language understanding.

[0102] For example, the external electronic device (302) can generate a prompt based on an audio signal and / or sensed data using the second multimodal model (360). For example, the external electronic device (302) can input or provide the generated prompt to a large language model (390). For example, the large language model (390) can generate response information based on the input of the generated prompt. For example, the external electronic device (302) can obtain response information for the audio signal using the large language model (390).

[0103] Figure 4a shows an example of learned actions of the first multimodal model.

[0104] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0105] According to one embodiment, operations 411 to 415 may be understood to be performed by a processor (e.g., at least one processor (300) of FIG. 3A) of an electronic device (e.g., electronic device (301) of FIG. 3A).

[0106] Referring to FIG. 4A, in operation 411, for example, an audio signal and sensing data may be input to a first multimodal model (e.g., the first multimodal model (350)). For example, the audio signal may represent a learning audio signal. For example, the sensing data may represent learning sensing data.

[0107] In operation 412, the electronic device (301) may include a feature extractor (or encoder). For example, the first multimodal model (350) may include the feature extractor. For example, the feature extractor may extract features of data of the audio signal from the audio signal and the sensing data. For example, the feature extractor may extract features of the sensing data from the sensing data. For example, the first multimodal model (350) may extract an embedding vector from the audio signal and the sensing data using the feature extractor.

[0108] In operation 413, the training of the first multimodal model (350) may be performed based on supervised learning and / or unsupervised learning. For example, the first multimodal model (350) may compare the similarity between the extracted embedding vector and the ground truth. For example, based on the result of the comparison, the first multimodal model (350) may obtain a loss function. For example, the first multimodal model (350) may be trained using the loss function. For example, the first multimodal model (350) may change the connection weights between nodes included in each of the layers (e.g., an input layer, one or more hidden layers, and an output layer) during training.

[0109] In operation 414, the first multimodal model (350) can adjust the probability of each modality for dropout. For example, the dropout may indicate partially excluding neurons in the neural network of the first multimodal model (350) from learning. For example, the first multimodal model (350) can avoid overfitting through the dropout.

[0110] For example, by inputting the embedding vector for the audio signal to the first multimodal model (350), the first multimodal model (350) can adjust the first probability for dropout of the audio signal. For example, by inputting the embedding vector for the sensing data to the first multimodal model (350), the first multimodal model (350) can adjust the second probability for dropout of the sensing data. For example, by inputting the embedding vector for the audio signal and the embedding vector for the sensing data to the first multimodal model (350), the first multimodal model (350) can adjust the third probability for dropout of the audio signal and the sensing data.

[0111] In operation 415, the first multimodal model (350) may perform dropout and fine tuning. For example, the first multimodal model (350) may perform dropout and fine tuning using the first probability, the second probability, and the third probability. For example, the first multimodal model (350) may change the connection weights between nodes included in each layer during tuning.

[0112] For example, although not illustrated in FIG. 4a, the learned motion of the first multimodal model (350) may correspond to the learned motion of the second multimodal model (360) of FIG. 3b. For example, the second multimodal model (360) of FIG. 3b may be learned with the motion illustrated in FIG. 4a.

[0113] Figure 4b shows an example of learned actions of the fourth multimodal model.

[0114] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0115] According to one embodiment, operations 421 to 425 may be understood to be performed by a processor (e.g., at least one processor (300) of FIG. 3A) of an electronic device (e.g., electronic device (301) of FIG. 3A).

[0116] Referring to FIG. 4B, in operation 421, the electronic device (301) may include a feature extractor (not shown). For example, a fourth multimodal model (e.g., the fourth multimodal model (380)) may include the feature extractor. For example, the fourth multimodal model (380) may extract features of the learning data from the learning data using the feature extractor. For example, the fourth multimodal model (380) may extract an embedding vector from the learning data using the feature extractor.

[0117] For example, the training data may be stored in a memory (e.g., memory (320)). For example, the training data may include an audio signal. For example, the feature extractor may extract features of the training data by applying a short time Fourier transform (STFT). For example, using the feature extractor, voice features may be extracted from the audio signal.

[0118] For example, in operation 422, the fourth multimodal model (380) can classify (or identify) the environment in which the training audio signal is acquired from the training audio signal. For example, the fourth multimodal model (380) can classify (or identify) the environment in which the training data is acquired from the embedding vector. For example, the fourth multimodal model (380) can classify (or identify) the environment using a third model (e.g., the third model (370)). For example, the third model (370) can classify the environment and provide data about the environment to the fourth multimodal model (380). For example, the fourth multimodal model (380) can use the data about the classified environment as a token.

[0119] In operation 423, the voice features extracted from the audio signal, the data about the classified environment, and the features of the learning sensing data may be input to a fourth multimodal model (380). For example, the voice features and the data about the classified environment may be input to the fourth multimodal model (380) after a data concatenation operation is performed.

[0120] In operation 424, the fourth multimodal model (380) may include an encoder. For example, the fourth multimodal model (380) may use the encoder to perform fusion of the speech features, the data about the environment, and the features of the learning sensing data. For example, the fusion may refer to combining different types of modalities into a single data. For example, as the fourth multimodal model (380) performs the fusion, the features of a single multimodal model may be acquired.

[0121] In operation 425, the fourth multimodal model (380) may perform noise cancellation on the audio signal by utilizing the characteristics of the single multimodal model. For example, the noise cancellation may be performed by utilizing an activation function (e.g., a sigmoid function). For example, the fourth multimodal model (380) may obtain a loss function based on the audio signal on which the noise cancellation was performed and the ground truth. For example, the fourth multimodal model (380) may be trained through the loss function. For example, the fourth multimodal model (380) may change the connection weights between nodes included in each of the layers (e.g., an input layer, one or more hidden layers, and an output layer) while being trained.

[0122] FIG. 5 illustrates examples of operations in which an electronic device transmits a signal that causes a wake-up of a second multimodal model. The operations of the electronic device (e.g., electronic device (301)) illustrated in FIG. 5 may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)).

[0123] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0124] Referring to FIG. 5, in operation 511, at least one processor (300) may detect an event that activates at least one sensor. For example, the event may be referenced as an input for driving at least one sensor (340). For example, at least one processor (300) may activate at least one sensor (340) by detecting the event.

[0125] For example, at least one sensor (340) may be in an inactive state before operation 511 is performed by at least one processor (300). For example, the image sensor may not be driven before at least one processor (300) performs operation 511. For example, activating at least one sensor (340) may include at least one processor (300) driving the image sensor.

[0126] As a non-limiting example, at least one sensor (340) may be running before operation 511 is executed by at least one processor (300). For example, activating at least one sensor (340) may include utilizing sensed data acquired via the at least one sensor (340) in connection with audio signal processing. For example, a heart rate sensor, a motion sensor, and / or an acceleration sensor may be running before operation 511 is executed. For example, activating at least one sensor (340) may include utilizing sensed data acquired via the at least one processor (300) via the heart rate sensor, the motion sensor, and / or the acceleration sensor in connection with audio signal processing. The above events are described and illustrated in more detail with reference to FIGS. 6A and 6B .

[0127] Figures 6a and 6b illustrate examples of events for driving at least one sensor.

[0128] Referring to FIG. 6A, the electronic device (301) may include an input element (610). For example, the input element (610) may be exposed through a portion of the housing of the electronic device (301). For example, the event may include at least one processor (300) receiving an input signal through the input element (610). For example, at least one processor (300) may detect the input signal received through the input element (610). For example, at least one processor (300) may activate at least one sensor (340) based on the detection.

[0129] For example, the input element (610) may be pressable. For example, the input element (610) may be a physical button. For example, the input element (610) may represent a pressable input button. For example, the event may include at least one processor (300) receiving a push input via the input element (610).

[0130] For example, the input element (610) may be rotatable. For example, the input element (610) may be rotatable relative to the housing of the electronic device (301). For example, the event may include at least one processor (300) receiving an input that rotates the input component via the input element (610).

[0131] For example, the input element (610) may include a touch sensor. For example, the event may include, but is not limited to, at least one processor (300) receiving a touch input through the input element (610).

[0132] For example, the electronic device (301) is not limited to an electronic device (301) configured in the shape of a watch. As a non-limiting example, the electronic device (301) may include a portable electronic device such as a smartphone, a tablet, a laptop computer, or a smartwatch.

[0133] Referring to FIG. 6B, an event for activating at least one sensor (340) may include detecting a reference motion. For example, based on the detection of the event, at least one sensor (340) in an inactive state may be driven. For example, the electronic device (301) may include a motion sensor and / or an acceleration sensor. For example, the motion sensor and / or the acceleration sensor may be in an activated state. For example, the motion sensor and / or the acceleration sensor may be in a state for low power consumption. For example, at least one processor (300) may detect the reference motion using the motion sensor and / or the acceleration sensor.

[0134] For example, the reference movement may include a movement of the electronic device (301) changing from state (620) to state (630). For example, state (620) may represent a state in which the direction of the front side of the electronic device (301) is not facing the user. For example, state (630) may represent a state in which the direction of the front side of the electronic device (301) is facing the user.

[0135] For example, the front side may be referred to as the front surface of the electronic device (301). For example, the front side may be referred to as an area that includes the display (311) of the electronic device (301).

[0136] For example, the reference movement may be related to the speed at which the electronic device (301) changes from state (620) to state (630). For example, the reference movement may be related to the difference between the direction of the front side of the electronic device (301) in state (620) and the direction of the front side of the electronic device (301) in state (630).

[0137] Referring back to FIG. 5, at operation 512, at least one processor (300) may acquire an audio signal via a microphone (330) based on detecting the event. For example, at least one processor (300) may acquire sensing data (e.g., an image (e.g., acquired via an image sensor), heart rate data (e.g., acquired via a heart rate sensor), and / or movement data (e.g., acquired via an acceleration sensor and / or a gyro sensor)) via at least one sensor (340) activated according to the event, based on detecting the event. For example, at least one processor (300) may acquire an image via an image sensor driven according to the event, based on detecting the event.

[0138] In operation 513, at least one processor (300) may identify whether the sensing data represents reference sensing data. For example, at least one processor (300) may identify whether the audio signal includes a reference voice command. For example, at least one processor (300) may identify the sensing data representing the reference sensing data of the user and the audio signal including the reference voice command using a first multimodal model (e.g., the first multimodal model (350)). For example, the reference sensing data may be preset by the user.

[0139] For example, at least one processor (300) may identify whether the image represents a reference gesture of the user. For example, at least one processor (300) may identify the image representing the reference gesture of the user and the audio signal including the reference voice command using a first multimodal model (350). For example, the reference gesture may be preset by the user. For example, the reference gesture is described and exemplified in more detail with reference to FIG. 7.

[0140] Figure 7 illustrates examples of reference gestures and reference voice commands recognized by the first multimodal model.

[0141] Referring to FIG. 7, in state (710), the user may indicate the reference gesture. For example, the reference gesture may indicate a gesture in which the user places the user's finger to the user's mouth.

[0142] For example, in state (710), the user may utter a reference voice command (720) while performing the reference gesture. For example, the reference voice command (720) may include a single syllable of speech. For example, the reference voice command (720) may include a two-syllable speech. For example, the reference voice command (720) may indicate 'shh'. However, the present invention is not limited thereto. The reference gesture and reference voice command (720) illustrated in FIG. 7 are merely examples.

[0143] Referring again to FIG. 5 , at operation 514, at least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the audio signal including the sensing data representing the reference sensing data and the reference voice command using the first multimodal model (350). For example, the signal may include a signal causing a wake-up of a second multimodal model (e.g., the second multimodal model (360)) within the external electronic device (302).

[0144] For example, at least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the image representing the reference gesture of the user and the audio signal including the reference voice command using the first multimodal model (350). For example, the signal may include a signal causing a wake-up of a second multimodal model (360) within the external electronic device (302).

[0145] For example, in operation 515, the external electronic device (302) may receive the signal via the communication circuit. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to wake up. For example, in response to receiving the signal, the external electronic device (302) may cause the activation of the second multimodal model (360).

[0146] For example, the electronic device (301) can cause the second multimodal model (360) within the external electronic device (302) to wake up in an environment where the volume of noise is relatively large by having at least one processor (300) perform the operations illustrated in FIG. 5. For example, the electronic device (301) can utilize a voice recognition function through the second multimodal model (360) in an environment where the volume of noise is relatively large.

[0147] For example, the electronic device (301) may cause the second multimodal model (360) within the external electronic device (302) to wake up in an environment where a user must whisper an utterance, by having at least one processor (300) perform the exemplary operations of FIG. 5 . For example, the electronic device (301) may utilize a voice recognition function through the second multimodal model (360) in an environment where a user must whisper an utterance. For example, the quality of the voice recognition function of the electronic device (301) may be enhanced.

[0148] FIGS. 8A and 8B illustrate examples of operations in which an electronic device transmits a signal that causes a wake-up of a second multimodal model. The operations of the electronic device (e.g., electronic device (301)) illustrated in FIGS. 8A and 8B may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0149] Referring to FIG. 8A, in operation 811, at least one processor (300) may obtain a first audio signal. For example, at least one processor (300) may obtain the first audio signal through a microphone (330).

[0150] For example, in operation 812, at least one processor (300) may identify the type of environment in which the electronic device (301) is located. For example, in operation 812, at least one processor (300) may identify the type of environment in which the electronic device (301) is located from the first audio signal. For example, at least one processor (300) may identify the type of environment using a third model (e.g., the third model (370)). As a non-limiting example, at least one processor (300) may identify the type of environment in which the electronic device (301) is located through sensing data acquired through at least one sensor (e.g., at least one sensor (340)).

[0151] For example, the type of the environment may include the first type and the second type. For example, for the first type and the second type, reference may be made to the descriptions of the third model (370) of FIG. 3B. For example, at least one processor (300) may identify an environment as the first type if the SNR of the environment in which the electronic device (301) is located is equal to or greater than a reference value. For example, at least one processor (300) may identify an environment as the second type if the SNR of the environment in which the electronic device (301) is located is less than a reference value.

[0152] For example, at least one processor (300) may execute operation 813 based on identifying that the type of the environment is the first type. For example, at least one processor (300) may execute operation 815 based on identifying that the type of the environment is the second type.

[0153] For example, at operation 813, at least one processor (300) may identify that the reference voice command is included in the first audio signal based on identifying that the type of the environment is the first type. For example, at least one processor (300) may identify that the reference voice command is included in the first audio signal using the first multimodal model (350). For example, at operation 813, at least one processor (300) may identify whether the reference voice command is included in the first audio signal.

[0154] For example, in operation 814, at least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) to cause a wake-up of the second multimodal model (360) according to the reference voice command included in the first audio signal.

[0155] For example, at operation 815, at least one processor (300) may activate at least one sensor (340) based on identifying that the type of the environment is the second type. For example, at least one processor (300) may drive at least one sensor (340) based on identifying that the type is the second type. For example, at least one sensor (340) may be in an inactive state prior to operation 815 being executed by at least one processor (300).

[0156] For example, the image sensor may not be running before at least one processor (300) executes operation 815. For example, activating at least one sensor (340) may include at least one processor (300) driving the image sensor.

[0157] As a non-limiting example, at least one sensor (340) may be running before operation 815 is executed by at least one processor (300). For example, activating at least one sensor (340) may include utilizing sensed data acquired via the at least one sensor (340) in connection with audio signal processing. For example, a heart rate sensor, a motion sensor, and / or an acceleration sensor may be running before operation 815 is executed. For example, activating at least one sensor (340) may include utilizing sensed data acquired via the at least one processor (300) via the heart rate sensor, the motion sensor, and / or the acceleration sensor in connection with audio signal processing.

[0158] For example, at least one processor (300) may display a screen including content notifying the activation of at least one sensor (340) through the display (311) before activating at least one sensor (340).

[0159] For example, in operation 816, at least one processor (300) may obtain the first sensing data (e.g., a first image (e.g., obtained via an image sensor), first heartbeat data (e.g., obtained via a heartbeat sensor), and / or first movement data (e.g., obtained via an acceleration sensor and / or a gyro sensor)) via at least one sensor (340). For example, at least one processor (300) may obtain the first image via an image sensor. For example, at least one processor (300) may obtain a second audio signal via a microphone (330).

[0160] For example, in operation 817, at least one processor (300) may identify reference sensing data and reference voice command using a first multimodal model (350). For example, at least one processor (300) may identify that the reference voice command is included in the second audio signal. For example, at least one processor (300) may identify whether the reference voice command is included in the second audio signal, in operation 817. For example, at least one processor (300) may identify that the reference sensing data is indicated in the first sensing data. For example, at least one processor (300) may identify whether the reference sensing data is indicated in the first sensing data. For example, at least one processor (300) may identify the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using a first multimodal model (350).

[0161] For example, at least one processor (300) may identify whether the reference gesture is included in the first image. For example, at least one processor (300) may identify the first image representing the reference gesture and the second audio signal including the reference voice command using the first multimodal model (350). For example, operation 817 may correspond to operation 513 of FIG. 5 .

[0162] For example, in operation 818, at least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, the signal may be transmitted to the external electronic device (302) via the communication circuit (310) based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using the first multimodal model (350). For example, the signal may represent a signal that causes a wake-up of the second multimodal model (360). For example, operation 818 may correspond to operation 514 of FIG. 5.

[0163] For example, in operation 819, the external electronic device (302) may receive the signal via the communication circuit. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to wake up. For example, in response to receiving the signal, the external electronic device (302) may cause the activation of the second multimodal model (360). For example, operation 819 may correspond to operation 515 of FIG. 5 .

[0164] Referring to FIG. 8B, in operation 821, at least one processor (300) may activate a timer. For example, at least one processor (300) may execute the timer. For example, operation 821 may be executed based on at least one processor (300) identifying in operation 812 of FIG. 8A that the type of the environment is the second type.

[0165] For example, at least one processor (300) may activate the timer based on identifying that the type of the environment is the second type. For example, the timer may be associated with at least one sensor (340). For example, the timer may be associated with the image sensor. For example, at least one processor (300) may activate at least one sensor (340) while the timer is activated. For example, at least one processor (300) may drive the image sensor while the timer is activated.

[0166] At operation 822, at least one processor (300) may activate at least one sensor (340) based on identifying that the type of the environment is the second type. For example, operation 822 may correspond to operation 815.

[0167] In operation 823, at least one processor (300) may obtain the first sensing data through at least one sensor (340). For example, at least one processor (300) may obtain a first image through an image sensor. For example, at least one processor (300) may obtain a second audio signal through a microphone (330). For example, operation 823 may correspond to operation 816.

[0168] In operation 824, at least one processor (300) may identify whether the reference voice command is included in the second audio signal. For example, at least one processor (300) may identify whether the reference gesture is included in the first sensing data. For example, at least one processor (300) may identify the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using a first multimodal model (350).

[0169] For example, at least one processor (300) may identify whether the reference gesture is included in the first image. For example, at least one processor (300) may identify the first image representing the reference gesture and the second audio signal including the reference voice command using the first multimodal model (350).

[0170] For example, at least one processor (300) may execute operations 825 and 827 based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command. For example, at least one processor (300) may execute operation 828 based on identifying the first sensing data not representing the reference sensing data and / or the second audio signal not including the reference voice command.

[0171] In operation 825, at least one processor (300) may transmit a signal to an external electronic device (302) via the communication circuit (310) based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using a first multimodal model (350). For example, at least one processor (300) may transmit the signal to the external electronic device (302) via the communication circuit (310) based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using a first multimodal model (350).

[0172] For example, the signal may be represented as a signal that causes the wake-up of the second multimodal model (360). For example, operation 825 may correspond to operation 818.

[0173] In operation 826, the external electronic device (302) may receive the signal via the communication circuit. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to wake up. For example, in response to receiving the signal, the external electronic device (302) may cause the second multimodal model (360) to be activated. For example, operation 826 may correspond to operation 819.

[0174] At operation 827, at least one processor (300) can maintain the activation of at least one sensor (340) based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, the at least one processor (300) can control the at least one sensor (340) to maintain the activation of the at least one sensor (340). For example, the at least one processor (300) can maintain operating the at least one sensor (340) independently of the expiration of the timer.

[0175] For example, at least one processor (300) can maintain the activation of at least one sensor (340) independently of the expiration of the timer based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, at least one processor (300) can maintain the operation of the image sensor independently of the expiration of the timer.

[0176] For example, at least one processor (300) may extend the remaining time of the timer based on identifying the first sensing data representing the reference sensing data and the second audio signal including the reference voice command using the first multimodal model (350). For example, as the remaining time of the timer is extended, the activation state of at least one sensor (340) may be maintained. For example, as the remaining time of the timer is extended, the operation of the image sensor may be maintained.

[0177] At step 828, at least one processor (300) may deactivate at least one sensor (340) based on identifying the first sensing data that does not represent the reference sensing data and / or the second audio signal that does not include the reference voice command. For example, the at least one processor (300) may control the at least one sensor (340) to deactivate the at least one sensor (340). For example, the at least one processor (300) may suspend an activated state of the at least one sensor (340). For example, the suspending may be executed in response to the expiration of the timer. For example, the suspending may include suspending operation of the image sensor. For example, the suspending may include suspending use of data acquired via a heart rate sensor, an acceleration sensor, and / or a gyro sensor for audio signal processing.

[0178] For example, at least one processor (300) may change the state of at least one sensor (340) from an activated state to a deactivated state based on identifying the first reference sensing data that does not represent the reference sensing data and / or the second audio signal that does not include the reference voice command. For example, the change may be performed in response to the expiration of the timer.

[0179] For example, at least one processor (300) may stop operation of the image sensor based on identifying the first image that does not express the reference gesture and / or the second audio signal that does not include the reference voice command. For example, the stopping may be performed in response to the expiration of the timer.

[0180] For example, the electronic device (301) can cause the wake-up of the second multimodal model (360) within the external electronic device (302) to be efficiently performed according to the environment in which the electronic device (301) is located by having at least one processor (300) perform the exemplary operations of FIGS. 8A to 8B.

[0181] For example, the electronic device (301) may cause a wake-up of the second multimodal model (360) within the external electronic device (302) in an environment where the noise volume is relatively large. For example, the electronic device (301) may use the second multimodal model (360) for a voice recognition function in an environment where the noise volume is relatively large.

[0182] For example, the electronic device (301) can cause the second multimodal model (360) within the external electronic device (302) to wake up in an environment where a user must whisper speech, by having at least one processor (300) perform the exemplary operations of FIGS. 8A and 8B . For example, the electronic device (301) can utilize the second multimodal model (360) for a speech recognition function in an environment where a user must whisper speech. For example, the quality of the speech recognition function of the electronic device (301) can be enhanced.

[0183] FIGS. 9A and 9B illustrate examples of operations in which an electronic device transmits data to an external electronic device and receives response information from the external electronic device to utilize a second multimodal model. The operations of the electronic device (e.g., electronic device (301)) illustrated in FIGS. 9A and 9B may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0184] Referring to FIG. 9A, operation 911 may represent an operation subsequent to operation 514 of FIG. 5. For example, operation 911 may represent an operation subsequent to operation 818 of FIG. 8A. For example, operation 911 may represent an operation subsequent to operation 827 of FIG. 8B. For example, operation 911 may be described as an operation after at least one processor (300) transmits, via the communication circuit (310), the signal causing the second multimodal model (360) to wake up.

[0185] For example, in operation 911, at least one processor (300) may obtain a third audio signal via a microphone (330). For example, at least one processor (300) may obtain second sensing data (e.g., a second image (e.g., obtained via an image sensor), second heartbeat data (e.g., obtained via a heartbeat sensor), and / or second movement data (e.g., obtained via an acceleration sensor and / or a gyro sensor)) via at least one sensor (340). For example, at least one processor (300) may obtain a second image via the image sensor.

[0186] For example, at least one processor (300) may obtain first data using the third audio signal. For example, the first data may represent data regarding an enhanced signal of the third audio signal. For example, the first data may represent data regarding an amplified signal of a user's voice signal within the third audio signal. For example, the first data may represent data regarding a noise-removed signal within the third audio signal. For example, the first data may represent data regarding feature extraction performed on the third audio signal for use in the second multimodal model (360).

[0187] For example, at least one processor (300) may acquire second data using the second sensing data. For example, the second data may represent data regarding a visual object within the second image. For example, the visual object may correspond to the lips of a user of the electronic device (301). For example, the second data may represent data related to the movement and / or shape of the visual object corresponding to the lips. However, this is not limited thereto.

[0188] For example, the electronic device (301) may include a face detection model (not shown). For example, the face detection model may identify a user's face from an image acquired through an image sensor. For example, the face detection model may identify an area containing another visual object corresponding to the user's face within the second image. For example, at least one processor (300) may obtain data regarding the area. For example, at least one processor (300) may use the data regarding the area to obtain second data regarding a visual object corresponding to the lips.

[0189] In operation 912, at least one processor (300) can transmit the first data and the second data to an external electronic device (302) via a communication circuit (310).

[0190] In operation 913, the external electronic device (302) may operate (or utilize) the second multimodal model (360). For example, the external electronic device (302) may receive the first data and the second data via a communication circuit. For example, the external electronic device (302) may provide (or input) the first data and the second data to the second multimodal model (360). For example, the second multimodal model (360) may be in a wake-up state.

[0191] For example, the external electronic device (302) can obtain data for input into the large language model (390) using the second multimodal model (360). For example, the data for input into the large language model (390) can be represented as a prompt. For example, the external electronic device (302) can generate a prompt for input into the large language model (390) by providing the first data and the second data to the second multimodal model (360).

[0192] In operation 914, the external electronic device (302) may operate the large language model (390). For example, the external electronic device (302) may obtain response information by inputting the prompt into the large language model. For example, the response information may be based on the first data and / or the second data. For example, the response information may represent a response to the third audio signal. For example, the response information may include a response to a visual object within the second image. The response information will be described and exemplified in more detail with reference to FIGS. 10A and 10B .

[0193] In operation 915, the external electronic device (302) can transmit the response information to the electronic device (301) via the communication circuit.

[0194] In operation 916, at least one processor (300) may perform or execute a function according to the response information. For example, at least one processor (300) may apply TTS (text to speech) to the response information. For example, at least one processor (300) may obtain an audio signal by applying TTS to the response information. For example, at least one processor (300) may output the audio signal obtained by applying TTS to the response information to the outside through a speaker (312).

[0195] For example, at least one processor (300) may display text within the response information through a display (311) by performing the function according to the response information. For example, the text within the response information may be based on the first data and / or the second data. For example, the text within the response information may be generated by a second multimodal model (360) based on the first data and / or the second data.

[0196] The above function according to the above response information will be described and illustrated in more detail with reference to FIGS. 10a and 10b.

[0197] Referring to FIG. 9B, in operation 921, at least one processor (300) may obtain the third audio signal via the microphone (330). For example, at least one processor (300) may obtain the second sensing data (e.g., a second image (e.g., obtained via an image sensor), second heartbeat data (e.g., obtained via a heartbeat sensor), and / or second movement data (e.g., obtained via an acceleration sensor and / or a gyro sensor)) via at least one sensor (340). For example, at least one processor (300) may obtain the second image via the image sensor. Operation 921 may correspond to operation 911. For example, operation 921 may represent a subsequent operation of operation 814 of FIG. 8A.

[0198] At operation 922, at least one processor (300) may identify the type of environment in which the electronic device (301) is located. For example, at operation 922, at least one processor (300) may identify the type of environment in which the electronic device (301) is located from the third audio signal. For example, at least one processor (300) may identify the type of environment using the third model (370). As a non-limiting example, at least one processor (300) may identify the type of environment using the second sensing data acquired through at least one sensor (340).

[0199] For example, at least one processor (300) may identify an environment in which an electronic device (301) is located as the first type of environment if the SNR of the environment is greater than or equal to a reference value. For example, at least one processor (300) may identify an environment in which an electronic device (301) is located as the second type of environment if the SNR of the environment is less than or equal to a reference value.

[0200] For example, at least one processor (300) may execute operation 923 based on identifying that the type of the environment is the first type. For example, at least one processor (300) may execute operation 924 and / or operation 925 based on identifying that the type of the environment is the second type.

[0201] In operation 923, at least one processor (300) may transmit data regarding the third audio signal and data regarding the second sensing data. For example, at least one processor (300) may obtain the second data using the second sensing data. For example, at least one processor (300) may transmit data regarding the third audio signal and the second data regarding the visual object within the second image.

[0202] For example, at least one processor (300) may obtain data for the third audio signal as the first data based on identifying that the type of the environment is the first type. For example, at least one processor (300) may transmit the first data and the second data to an external electronic device (302) via a communication circuit (310). For example, operation 923 may correspond to operation 912. For example, after at least one processor (300) executes operation 923, the external electronic device (302) may execute operations 913 to 915. For example, after the external electronic device (302) executes operations 913 to 915, the at least one processor (300) may execute operation 916.

[0203] At step 924, at least one processor (300) may perform noise cancellation on the third audio signal. For example, at least one processor (300) may perform noise cancellation on the third audio signal based on identifying that the type of the environment is the second type. For example, at least one processor (300) may perform noise cancellation using the fourth multimodal model (380).

[0204] For example, at least one processor (300) can perform the noise canceling by providing (or inputting) the second data, the data for the second type, and / or the data for the third audio signal to the fourth multimodal model (380). For example, at least one processor (300) can obtain a fourth audio signal from which noise is removed from the third audio signal by providing the second data, the data for the second type, and / or the data for the third audio signal to the fourth multimodal model (380). For example, the fourth audio signal can be generated by performing noise canceling on the third audio signal by the fourth multimodal model (380).

[0205] For example, the fourth audio signal may represent an audio signal in which noise within the third audio signal is reduced. For example, the fourth audio signal may represent an audio signal in which a voice signal within the third audio signal is enhanced.

[0206] For example, at least one processor (300) may obtain data for the fourth audio signal on which noise cancellation is performed on the third audio signal. For example, at least one processor (300) may obtain data for the fourth audio signal on which noise cancellation is performed on the third audio signal as the first data. For example, the data for the fourth audio signal may represent data on which feature extraction is performed on the fourth audio signal for use by the second multimodal model (360).

[0207] For example, the SNR of the fourth audio signal may be higher than the SNR of the third audio signal. For example, since the SNR of the fourth audio signal is higher than the SNR of the third audio signal, the ASR function of the second multimodal model (360) may be enhanced when data for the fourth audio signal is input to the second multimodal model (360) compared to when data for the third audio signal is input to the second multimodal model (360).

[0208] In operation 925, at least one processor (300) may transmit the first data and the second data to the external electronic device (302) via the communication circuit (310). For example, operation 925 may correspond to operation 912. For example, after the at least one processor (300) executes operation 925, the external electronic device (302) may execute operations 913 to 915. For example, after the external electronic device (302) executes operations 913 to 915, the at least one processor (300) may execute operation 916.

[0209] For example, by having at least one processor (300) perform the exemplary operations of FIGS. 9A to 9B, the electronic device (301) can determine whether to perform noise canceling on an audio signal acquired through a microphone (330) according to the environment.

[0210] For example, the electronic device (301) can efficiently utilize the second multimodal model (360) within the external electronic device (302) depending on the environment in which the electronic device (301) is located. For example, the electronic device (301) can reduce the current consumed for utilizing the second multimodal model (360) by determining whether to perform noise canceling depending on the environment.

[0211] For example, the electronic device (301) can enhance the quality of the voice recognition function by performing noise cancellation on the acquired audio signal using the fourth multimodal model (380). For example, the electronic device (301) can enhance the quality of the voice recognition function by transmitting sensing data acquired through at least one sensor (340) to an external electronic device (302).

[0212] Figures 10a and 10b illustrate examples of electronic devices that perform functions according to response information.

[0213] Referring to FIG. 10A, a state in which a user utters a voice signal to an electronic device (301) may be described. For example, in state (1010), at least one processor (300) may obtain the second sensing data through at least one sensor (340). For example, at least one processor (300) may obtain the third audio signal through a microphone (330).

[0214] For example, in state (1020), at least one processor (300) may perform a function according to response information in operation 916. For example, in state (1020), for example, at least one processor (300) may perform the function according to the response information, thereby displaying an emoji graphical object (1021) appearing in the response information among emoji graphical objects available in the electronic device (301) through the display (311).

[0215] For example, the emoji graphical object (1021) may be selected by the second multimodal model (360). For example, the emoji graphical object (1021) may be selected by the second multimodal model (360) based on the first data and / or the second data. For example, the emoji graphical object (1021) may be pre-stored in the memory (320) in relation to the first data and / or the second data. For example, the emoji graphical object (1021) may be pre-stored in the external electronic device (302) in relation to the first data and / or the second data.

[0216] The emoji graphical object (1021) illustrated in FIG. 10A is merely exemplary. The emoji graphical object (1021) may include other emoji graphical objects (1021) than the city of FIG. 10A.

[0217] Referring to FIG. 10b, a state in which a user utters a voice signal to an electronic device (301) may be described. For example, in state (1030), at least one processor (300) may perform the operations exemplified in FIGS. 9a and 9b and perform a function according to response information. For example, at least one processor (300) may perform the function according to the response information, thereby displaying text (1031) indicated by the response information among texts pre-stored in the electronic device (301) and / or the external electronic device (302) through the display (311).

[0218] For example, text (1031) may be selected by the second multimodal model (360). For example, text (1031) may be selected by the second multimodal model (360) based on the first data and / or the second data. For example, text (1031) may be pre-stored in the memory (320) in relation to the first data and / or the second data. For example, text (1031) may be pre-stored in the external electronic device (302) in relation to the first data and / or the second data.

[0219] The text (1031) shown in FIG. 10b is merely exemplary. The text (1031) may include text (1031) different from that shown in FIG. 10b.

[0220] FIG. 11 illustrates an example of an electronic device performing a call connection with an external electronic device. The operations of the electronic device (e.g., electronic device (301)) illustrated in FIG. 11 may be executed, performed, or controlled by at least one processor (e.g., at least one processor (300)). In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0221] For example, in operation 1101, at least one processor (300) may detect an event for a call connection with an external electronic device. For example, the event may include receiving a touch input through a display (e.g., display (311)). For example, the event may include receiving a touch input for an executable object displayed through the display (311). For example, the event may include receiving a scroll input for an executable object displayed through the display (311). However, the present invention is not limited thereto. For example, the event may include an event exemplified in FIG. 6A.

[0222] For example, in operation 1102, at least one processor (300) may obtain a fifth audio signal via a microphone (e.g., microphone (330)).

[0223] For example, in operation 1103, at least one processor (300) may identify the type of environment in which the electronic device (301) is located. For example, at least one processor (300) may identify the type of environment in which the electronic device (301) is located from the fifth audio signal. For example, at least one processor (300) may identify the type of environment using a third model (e.g., the third model (370)). As a non-limiting example, at least one processor (300) may identify the type of environment in which the electronic device (301) is located through sensing data acquired through at least one sensor (e.g., at least one sensor (340)).

[0224] For example, the type of the environment may include the first type and the second type. For example, at least one processor (300) may identify the environment as the first type if the SNR of the environment in which the electronic device (301) is located is equal to or greater than a reference value. For example, at least one processor (300) may identify the environment as the second type if the SNR of the environment in which the electronic device (301) is located is less than a reference value.

[0225] For example, at least one processor (300) may execute operation 1104 based on identifying that the type of the environment is the first type. For example, at least one processor (300) may execute operation 1106 based on identifying that the type of the environment is the second type. For example, operation 1103 may correspond to operation 812 of FIG. 8A. For example, operation 1103 may correspond to operations 922 of FIG. 9B.

[0226] For example, in operation 1104, at least one processor (300) may obtain a sixth audio signal via a microphone (330). For example, the sixth audio signal may include a user's voice.

[0227] For example, in operation 1105, at least one processor (300) may transmit the sixth audio signal to an external electronic device via a communication circuit (310). For example, the external electronic device may include a server and a base station. For example, the external electronic device may include an electronic device of another user. For example, in operation 1105, the electronic device (301) may indicate a state in which it is in a call with the external electronic device.

[0228] For example, in operation 1106, at least one processor (300) may activate at least one sensor (340) based on identifying that the type of the environment is the second type. For example, at least one processor (300) may drive at least one sensor (340) based on identifying that the type is the second type. For example, at least one sensor (340) may be in a disabled state prior to operation 1106 being executed by the at least one processor (300). For example, operation 1106 may correspond to operation 815 of FIG. 8A.

[0229] For example, in operation 1107, at least one processor (300) may obtain third sensing data (e.g., a third image (e.g., obtained via an image sensor), third heartbeat data (e.g., obtained via a heartbeat sensor), and / or third movement data (e.g., obtained via an acceleration sensor and / or a gyro sensor)) via at least one sensor (340). For example, at least one processor (300) may obtain the third image via an image sensor. For example, at least one processor (300) may obtain a seventh audio signal via a microphone (330). For example, operation 1107 may correspond to operation 816 of FIG. 8A.

[0230] For example, in operation 1108, at least one processor (300) may provide the seventh audio signal and the third sensing data to a first multimodal model (e.g., the first multimodal model (350)). For example, at least one processor (300) may obtain response information for the seventh audio signal and the third sensing data. For example, the response information may represent a response to the seventh audio signal. For example, the response information may include a response to a user's utterance within the seventh audio signal. For example, the response information may represent a response to the third sensing data.

[0231] For example, in operation 1109, at least one processor (300) may perform a function according to response information for the seventh audio signal and the third sensing data. For example, the response information may cause the at least one processor (300) to display an emoji graphical object (e.g., an emoji graphical object (1021)) corresponding to the seventh audio signal through the display (311). For example, the emoji graphical object (1021) may be selected by the first multimodal model (350). For example, the emoji graphical object (1021) may be generated by the first multimodal model (350) based on the seventh audio signal and / or the third sensing data.

[0232] For example, the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310). For example, the external electronic device receiving the signal may display an emoji graphical object (1021) corresponding to the seventh audio signal through a display of the external electronic device.

[0233] For example, the response information may cause at least one processor (300) to display an emoji graphical object (1021) corresponding to the third sensing data through the display (311). For example, the third sensing data may include data related to a visual object corresponding to the user's lips in an image acquired through the image sensor.

[0234] For example, the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310). For example, the external electronic device receiving the signal may display an emoji graphical object (1021) corresponding to the seventh audio signal through a display of the external electronic device.

[0235] For example, the response information may cause at least one processor (300) to display text (e.g., text (1031)) corresponding to the seventh audio signal via the display (311). For example, the text (1031) may be selected by the first multimodal model (350).

[0236] For example, the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310). For example, the external electronic device receiving the signal may display text (1031) corresponding to the seventh audio signal through a display of the external electronic device.

[0237] For example, text (1031) may be selected by the first multimodal model (350). For example, text (1031) may be generated by the first multimodal model (350) based on the seventh audio signal and / or the third sensing data.

[0238] For example, the response information may cause at least one processor (300) to display text (1031) corresponding to third sensing data through the display (311). For example, the third sensing data may include data related to a visual object corresponding to the user's lips in an image acquired through the image sensor.

[0239] For example, the response information may cause at least one processor (300) to transmit a signal to an external electronic device via a communication circuit (310). For example, the external electronic device receiving the signal may display text (1031) corresponding to the seventh audio signal through a display of the external electronic device.

[0240] FIG. 12 is a block diagram of an electronic device (1201) within a network environment (1200) according to various embodiments. For example, the electronic device (1201) may include an electronic device (301).

[0241] Referring to FIG. 12, in a network environment (1200), an electronic device (1201) may communicate with an electronic device (1202) via a first network (1298) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1204) or a server (1208) via a second network (1299) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1201) may communicate with the electronic device (1204) via the server (1208). According to one embodiment, the electronic device (1201) may include a processor (1220), a memory (1230), an input module (1250), an audio output module (1255), a display module (1260), an audio module (1270), a sensor module (1276), an interface (1277), a connection terminal (1278), a haptic module (1279), a camera module (1280), a power management module (1288), a battery (1289), a communication module (1290), a subscriber identification module (1296), or an antenna module (1297). In some embodiments, the electronic device (1201) may omit at least one of these components (e.g., the connection terminal (1278)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1276), camera module (1280), or antenna module (1297)) may be integrated into a single component (e.g., display module (1260)).

[0242] The processor (1220) may control at least one other component (e.g., hardware or software component) of the electronic device (1201) connected to the processor (1220) by executing, for example, software (e.g., program (1240)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1220) may store commands or data received from other components (e.g., sensor module (1276) or communication module (1290)) in volatile memory (1232), process the commands or data stored in volatile memory (1232), and store result data in non-volatile memory (1234). According to one embodiment, the processor (1220) may include a main processor (1221) (e.g., a central processing unit or an application processor) or an auxiliary processor (1223) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1221). For example, when the electronic device (1201) includes the main processor (1221) and the auxiliary processor (1223), the auxiliary processor (1223) may be configured to use less power than the main processor (1221) or to be specialized for a given function. The auxiliary processor (1223) may be implemented separately from the main processor (1221) or as a part thereof.

[0243] The auxiliary processor (1223) may control at least a portion of functions or states associated with at least one component (e.g., a display module (1260), a sensor module (1276), or a communication module (1290)) of the electronic device (1201), for example, on behalf of the main processor (1221) while the main processor (1221) is in an inactive (e.g., sleep) state, or together with the main processor (1221) while the main processor (1221) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1223) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1280) or a communication module (1290)). In one embodiment, the auxiliary processor (1223) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1201) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1208)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0244] The memory (1230) can store various data used by at least one component (e.g., the processor (1220) or the sensor module (1276)) of the electronic device (1201). The data can include, for example, software (e.g., the program (1240)) and input data or output data for commands related thereto. The memory (1230) can include a volatile memory (1232) or a non-volatile memory (1234).

[0245] The program (1240) may be stored as software in memory (1230) and may include, for example, an operating system (1242), middleware (1244), or an application (1246).

[0246] The input module (1250) can receive commands or data to be used in a component of the electronic device (1201) (e.g., a processor (1220)) from an external source (e.g., a user) of the electronic device (1201). The input module (1250) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0247] The audio output module (1255) can output audio signals to the outside of the electronic device (1201). The audio output module (1255) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0248] The display module (1260) can visually provide information to an external party (e.g., a user) of the electronic device (1201). The display module (1260) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1260) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0249] The audio module (1270) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1270) can acquire sound through the input module (1250), output sound through the sound output module (1255), or an external electronic device (e.g., electronic device (1202)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1201).

[0250] The sensor module (1276) can detect the operating status (e.g., power or temperature) of the electronic device (1201) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1276) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0251] The interface (1277) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1201) with an external electronic device (e.g., the electronic device (1202)). In one embodiment, the interface (1277) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0252] The connection terminal (1278) may include a connector through which the electronic device (1201) may be physically connected to an external electronic device (e.g., the electronic device (1202)). In one embodiment, the connection terminal (1278) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0253] The haptic module (1279) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1279) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0254] The camera module (1280) can capture still images and videos. According to one embodiment, the camera module (1280) may include one or more lenses, image sensors, image signal processors, or flashes.

[0255] The power management module (1288) can manage the power supplied to the electronic device (1201). According to one embodiment, the power management module (1288) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0256] A battery (1289) may power at least one component of the electronic device (1201). In one embodiment, the battery (1289) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0257] The communication module (1290) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1201) and an external electronic device (e.g., electronic device (1202), electronic device (1204), or server (1208)), and the performance of communication through the established communication channel. The communication module (1290) may operate independently from the processor (1220) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1290) may include a wireless communication module (1292) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1294) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1204) via a first network (1298) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1299) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1292) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1296) to verify or authenticate the electronic device (1201) within a communication network such as the first network (1298) or the second network (1299).

[0258] The wireless communication module (1292) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1292) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1292) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1292) may support various requirements specified in the electronic device (1201), an external electronic device (e.g., the electronic device (1204)), or a network system (e.g., the second network (1299)). According to one embodiment, the wireless communication module (1292) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0259] The antenna module (1297) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1297) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1297) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1298) or the second network (1299), may be selected from the plurality of antennas, for example, by the communication module (1290). A signal or power may be transmitted or received between the communication module (1290) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1297).

[0260] According to various embodiments, the antenna module (1297) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0261] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0262] According to one embodiment, commands or data may be transmitted or received between the electronic device (1201) and an external electronic device (1204) via a server (1208) connected to a second network (1299). Each of the external electronic devices (1202 or 1204) may be the same or a different type of device as the electronic device (1201). According to one embodiment, all or part of the operations executed in the electronic device (1201) may be executed in one or more of the external electronic devices (1202, 1204, or 1208). For example, when the electronic device (1201) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1201) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1201). The electronic device (1201) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1201) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1204) may include an Internet of Things (IoT) device. The server (1208) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1204) or server (1208) may be included within the second network (1299). The electronic device (1201) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.

[0263] In embodiments of the present disclosure, an electronic device (e.g., electronic device (301)) may be an electronic device for displaying an image in a virtual space. For example, an electronic device (e.g., electronic device (301) of FIG. 1) for displaying an image in a virtual space may be a wearable device. Hereinafter, descriptions may be given of a wearable device (e.g., wearable device (1301) to be described below) as an example of an electronic device. A wearable device may include a head-mounted display (HMD) that is wearable on a user's head. For example, a wearable device may be referred to as a head-mounted electronic device. A wearable device may be referred to as a head-mounted device (HMD), a headgear electronic device, a glasses-type electronic device, a video see-through (VST) device, an extended reality (XR) device, a virtual reality (VR) device, and / or an augmented reality (AR) device. Although the external appearance of the wearable device in the form of glasses is shown, the embodiment is not limited thereto. An example of a hardware configuration included in a wearable device is exemplarily described with reference to FIG. 16. An example of a structure of a wearable device that can be worn on a user's head is described with reference to FIGS. 13A, 13B, 13C, 14A, 14B, and / or 15. A wearable device may be referred to as an electronic device. For example, a wearable device may be combined with an accessory (e.g., a strap) to attach to a user's head to form an HMD.

[0264] In one embodiment, a wearable device may perform functions related to augmented reality (AR) and / or mixed reality (MR). For example, while a user wears the wearable device, the wearable device may include at least one lens positioned adjacent to the user's eyes. The wearable device may combine ambient light passing through the lens with light emitted from a display of the wearable device. A display area of ​​the display may be formed within the lens through which the ambient light passes. Because the wearable device combines the ambient light and the light emitted from the display, the user may see an image that is a mixture of a real object perceived by the ambient light and a virtual object formed by the light emitted from the display. The augmented reality, mixed reality, and / or virtual reality described above may be referred to as extended reality (XR).

[0265] In one embodiment, a wearable device may perform functions related to video see-through (VST) and / or virtual reality (VR). For example, when a user wears the wearable device, the wearable device may include a housing that covers the user's eyes. The wearable device may include a display disposed on a first side of the housing facing the eyes. The wearable device may include a camera disposed on a second side opposite the first side. Using the camera, the wearable device may acquire images and / or videos representing ambient light. The wearable device may output the images and / or videos within the display disposed on the first side, thereby allowing the user to perceive the ambient light through the display. A displaying area (or displaying region) (or active area (or active region)) of the display disposed on the first side may be formed by one or more pixels included in the display. The wearable device can synthesize a virtual object into an image and / or video output through the display, thereby allowing the user to recognize the virtual object together with a real object recognized by ambient light.

[0266] In one embodiment, a wearable device can identify or recognize a position and / or direction or orientation of the wearable device based on an image (and / or video) obtained or acquired using a camera. The wearable device can obtain information about the external space using one or more cameras and / or one or more sensors. The information can include a geographic location (e.g., global positioning system (GPS) coordinates) of the external space identified from one or more sensors. The information can include images and / or videos of the external space identified from one or more cameras. The wearable device can perform object recognition on the images and / or videos to identify external objects included in the external space from the images and / or videos.

[0267] Hereinafter, an example of a hardware configuration of a wearable device is described with reference to FIGS. 13a, 13b, 13c, 14a, 14b, 15, and / or 16.

[0268] FIG. 13A illustrates an example of a perspective view of a wearable device. FIG. 13B illustrates an example of one or more hardware components arranged within the wearable device. FIG. 13C illustrates an example of a wearable device according to an embodiment. The wearable device (1301) may have the form of glasses that can be worn on a body part (e.g., head) of a user. The wearable device (1301) of FIGS. 13A to 13C may be an example of the electronic device (301) of FIG. 3A. The wearable device (1301) may include a head-mounted display (HMD). For example, the housing of the wearable device (1301) may include a flexible material, such as rubber and / or silicone, that is configured to fit closely to a portion of the user's head (e.g., a portion of the face surrounding both eyes). For example, the housing of the wearable device (1301) may include one or more straps capable of being twined around the user's head, and / or one or more temples attachable to the ears of the head.

[0269] Referring to FIG. 13A, according to one embodiment, a wearable device (1301) may include at least one display (1350) and a frame (1300) supporting at least one display (1350). For example, the at least one display (1350) may be an example of the display (311) of FIG. 3A.

[0270] According to one embodiment, a wearable device (1301) can be worn on a part of a user's body. The wearable device (1301) can provide augmented reality (AR), virtual reality (VR), or mixed reality (MR) that combines augmented reality and virtual reality to a user wearing the wearable device (1301). For example, the wearable device (1301) can display a virtual reality image provided from at least one optical device (1382, 1384) of FIG. 13B on at least one display (1350) in response to a user's designated gesture acquired through the motion recognition cameras (1360-2, 1360-3) of FIG. 13B.

[0271] According to one embodiment, at least one display (1350) may provide visual information to a user. For example, at least one display (1350) may include a transparent or translucent lens. At least one display (1350) may include a first display (1350-1) and / or a second display (1350-2) spaced apart from the first display (1350-1). For example, the first display (1350-1) and the second display (1350-2) may be positioned at positions corresponding to the user's left and right eyes, respectively.

[0272] Referring to FIG. 13B, at least one display (1350) can provide visual information transmitted from external light to a user through a lens included in at least one display (1350), and other visual information distinct from the visual information. The lens can be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens. For example, at least one display (1350) can include a first surface (1331) and a second surface (1332) opposite to the first surface (1331). A display area can be formed on the second surface (1332) of at least one display (1350). When a user wears the wearable device (1301), external light can be transmitted to the user by being incident on the first surface (1331) and transmitted through the second surface (1332). As another example, at least one display (1350) can display an augmented reality image combined with a virtual reality image provided from at least one optical device (1382, 1384) on a real screen transmitted through external light, in a display area formed on the second surface (1332).

[0273] In one embodiment, at least one display (1350) may include at least one waveguide (1333, 1334) that diffracts light emitted from at least one optical device (1382, 1384) and transmits the diffracted light to a user. The at least one waveguide (1333, 1334) may be formed based on at least one of glass, plastic, or polymer. A nano-pattern may be formed on at least a portion of the exterior or interior of the at least one waveguide (1333, 1334). The nano-pattern may be formed based on a grating structure having a polygonal and / or curved shape. Light incident on one end of the at least one waveguide (1333, 1334) may be propagated to the other end of the at least one waveguide (1333, 1334) by the nano-pattern. At least one waveguide (1333, 1334) may include at least one diffractive element (e.g., a diffractive optical element (DOE), a holographic optical element (HOE)), or at least one reflective element (e.g., a reflective mirror). For example, at least one waveguide (1333, 1334) may be arranged within the wearable device (1301) to guide a screen displayed by at least one display (1350) to the user's eyes. For example, the screen may be transmitted to the user's eyes based on total internal reflection (TIR) ​​occurring within the at least one waveguide (1333, 1334).

[0274] The wearable device (1301) can analyze an object included in a real image collected through a shooting camera (1360-4), combine a virtual object corresponding to an object to be provided with augmented reality among the analyzed objects, and display the virtual object on at least one display (1350). The virtual object can include at least one of text and an image regarding various information related to the object included in the real image. The wearable device (1301) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the wearable device (1301) can perform spatial recognition (e.g., simultaneous localization and mapping (SLAM)) using the multi-camera and / or time-of-flight (ToF). A user wearing the wearable device (1301) can view an image displayed on at least one display (1350).

[0275] According to one embodiment, the frame (1300) may be configured as a physical structure that allows the wearable device (1301) to be worn on the user's body. According to one embodiment, the frame (1300) may be configured so that, when the user wears the wearable device (1301), the first display (1350-1) and the second display (1350-2) can be positioned corresponding to the user's left and right eyes. The frame (1300) may support at least one display (1350). For example, the frame (1300) may support the first display (1350-1) and the second display (1350-2) to be positioned corresponding to the user's left and right eyes.

[0276] Referring to FIG. 13A, the frame (1300) may include a region (1320) that at least partially contacts a portion of the user's body when the user wears the wearable device (1301). For example, the region (1320) of the frame (1300) that contacts a portion of the user's body may include a region that contacts a portion of the user's nose, a portion of the user's ear, and a portion of the side of the user's face that the wearable device (1301) comes into contact with. According to one embodiment, the frame (1300) may include a nose pad (1310) that contacts a portion of the user's body. When the wearable device (1301) is worn by the user, the nose pad (1310) may contact a portion of the user's nose. The frame (1300) may include a first temple (1304) and a second temple (1305) that contact another part of the user's body that is distinct from the part of the user's body.

[0277] For example, the frame (1300) may include a first rim (1302-1) that surrounds at least a portion of the first display (1350-1), a second rim (1302-2) that surrounds at least a portion of the second display (1350-2), a bridge (1303) that is disposed between the first rim (1302-1) and the second rim (1302-2), a first pad (1311) that is disposed along a portion of the edge of the first rim (1302-1) from one end of the bridge (1303), a second pad (1312) that is disposed along a portion of the edge of the second rim (1302-2) from the other end of the bridge (1303), a first temple (1304) that extends from the first rim (1302-1) and is fixed to a portion of the wearer's ear, and a second temple (1304) that extends from the second rim (1302-2) and is fixed to the ear opposite the ear. A second temple (1305) may be fixed to a portion. The first pad (1311) and the second pad (1312) may be in contact with a portion of the user's nose, and the first temple (1304) and the second temple (1305) may be in contact with a portion of the user's face and a portion of the user's ear. The temples (1304, 1305) may be rotatably connected to the rim via the hinge units (1306, 1307) of FIG. 13B. The first temple (1304) may be rotatably connected to the first rim (1302-1) via the first hinge unit (1306) disposed between the first rim (1302-1) and the first temple (1304). The second temple (1305) can be rotatably connected to the second rim (1302-2) via a second hinge unit (1307) disposed between the second rim (1302-2) and the second temple (1305). In one embodiment, the wearable device (1301) can identify an external object (e.g., a user's fingertip) touching the frame (1300) and / or a gesture performed by the external object by using a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of a surface of the frame (1300).

[0278] According to one embodiment, the wearable device (1301) may include hardwares that perform various functions (e.g., hardwares to be described later based on the block diagram of FIG. 16). For example, the hardwares may include a battery module (1370), an antenna module (1375), at least one optical device (1382, 1384), speakers (e.g., speakers 1355-1, 1355-2), a microphone (e.g., microphones 1365-1, 1365-2, 1365-3), a light-emitting module (not shown), and / or a printed circuit board (PCB) (1390) (e.g., a printed circuit board). The various hardwares may be arranged within the frame (1300). For example, the microphones (1365-1, 1365-2, 1365-3) may be examples of the microphone (330) of FIG. 3A.

[0279] According to one embodiment, microphones (e.g., microphones 1365-1, 1365-2, 1365-3) of a wearable device (1301) may be disposed on at least a portion of a frame (1300) to acquire sound signals. A first microphone (1365-1) disposed on a bridge (1303), a second microphone (1365-2) disposed on a second rim (1302-2), and a third microphone (1365-3) disposed on the first rim (1302-1) are illustrated in FIG. 13B , but the number and arrangement of the microphones (1365) are not limited to the embodiment of FIG. 13B . When the number of microphones (1365) included in the wearable device (1301) is two or more, the wearable device (1301) can identify the direction of a sound signal by using a plurality of microphones arranged on different parts of the frame (1300).

[0280] According to one embodiment, at least one optical device (1382, 1384) may project a virtual object onto at least one display (1350) to provide various image information to a user. For example, at least one optical device (1382, 1384) may be a projector. At least one optical device (1382, 1384) may be disposed adjacent to at least one display (1350) or may be included within at least one display (1350) as a part of at least one display (1350). According to one embodiment, the wearable device (1301) may include a first optical device (1382) corresponding to a first display (1350-1) and a second optical device (1384) corresponding to a second display (1350-2). For example, at least one optical device (1382, 1384) may include a first optical device (1382) disposed at an edge of a first display (1350-1) and a second optical device (1384) disposed at an edge of a second display (1350-2). The first optical device (1382) may transmit light to a first waveguide (1333) disposed on the first display (1350-1), and the second optical device (1384) may transmit light to a second waveguide (1334) disposed on the second display (1350-2).

[0281] In one embodiment, the camera (1360) may include a recording camera (1360-4), an eye tracking camera (ET CAM) (1360-1), and / or a motion recognition camera (1360-2, 1360-3). The recording camera (1360-4), the eye tracking camera (1360-1), and the motion recognition cameras (1360-2, 1360-3) may be positioned at different locations on the frame (1300) and may perform different functions. The eye tracking camera (1360-1) may output data indicating the position or gaze of the eyes of a user wearing the wearable device (1301). For example, the wearable device (1301) may detect the gaze from an image including the user's pupils obtained through the eye tracking camera (1360-1). The wearable device (1301) can identify an object (e.g., a real object and / or a virtual object) focused on by the user using the user's gaze acquired through the gaze tracking camera (1360-1). The wearable device (1301) that has identified the focused object can execute a function (e.g., gaze interaction) for interaction between the user and the focused object. The wearable device (1301) can express a part corresponding to the eye of an avatar representing the user in a virtual space using the user's gaze acquired through the gaze tracking camera (1360-1). The wearable device (1301) can render an image (or screen) displayed on at least one display (1350) based on the position of the user's eyes. For example, the visual quality of a first area related to the gaze within the image and the visual quality (e.g., resolution, brightness, saturation, grayscale, PPI (pixels per inch)) of a second area distinguished from the first area may be different from each other. In this disclosure, the term “resolution” is used to refer to the density of pixels of an image and / or display (1350).The density and / or resolution of pixels can be measured or parameterized based on units of PPI and / or dpi (dots per inch). The wearable device (1301) can obtain an image having a visual quality of a first area matching the user's gaze and a visual quality of a second area using foveated rendering. For example, if the wearable device (1301) supports an iris recognition function, user authentication can be performed based on iris information obtained using the gaze tracking camera (1360-1). Although an example in which the gaze tracking camera (1360-1) is positioned toward the user's right eye is illustrated in FIG. 13B, the embodiment is not limited thereto, and the gaze tracking camera (1360-1) can be positioned solely toward the user's left eye, or toward both eyes.

[0282] In one embodiment, the capturing camera (1360-4) can capture an actual image or background to be aligned with a virtual image to implement augmented reality or mixed reality content. The capturing camera (1360-4) can be used to obtain a high-resolution image based on HR (high resolution) or PV (photo video). The capturing camera (1360-4) can capture an image of a specific object existing at a location viewed by the user and provide the image to at least one display (1350). The at least one display (1350) can display a single image in which information about an actual image or background including an image of the specific object obtained using the capturing camera (1360-4) and a virtual image provided through at least one optical device (1382, 1384) are superimposed. The wearable device (1301) can compensate for depth information (e.g., the distance between the wearable device (1301) and an external object acquired through a depth sensor) using an image acquired through the capture camera (1360-4). The wearable device (1301) can perform object recognition using an image acquired through the capture camera (1360-4). The wearable device (1301) can perform a function of focusing on an object (or subject) in an image (e.g., auto focus (AF)) and / or an optical image stabilization (OIS) function (e.g., anti-shake function) using the capture camera (1360-4). The wearable device (1301) can perform a pass-through function to display an image acquired through the capture camera (1360-4) by overlapping at least a portion of a screen representing a virtual space on at least one display (1350). In one embodiment, the shooting camera (1360-4) may be positioned on a bridge (1303) positioned between the first rim (1302-1) and the second rim (1302-2).

[0283] The gaze tracking camera (1360-1) can implement more realistic augmented reality by tracking the gaze of a user wearing a wearable device (1301) and matching the user's gaze with visual information provided to at least one display (1350). For example, when the wearable device (1301) looks straight ahead, the wearable device (1301) can naturally display environmental information related to the user's front at a location where the user is located on at least one display (1350). The gaze tracking camera (1360-1) can be configured to capture an image of the user's pupil to determine the user's gaze. For example, the gaze tracking camera (1360-1) can receive gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In one embodiment, the gaze tracking camera (1360-1) can be positioned at positions corresponding to the user's left and right eyes. For example, the gaze tracking camera (1360-1) may be positioned within the first rim (1302-1) and / or the second rim (1302-2) to face the direction in which the user wearing the wearable device (1301) is positioned.

[0284] The gesture recognition cameras (1360-2, 1360-3) can recognize the movement of the user's entire body, such as the user's torso, hands, or face, or a part of the body, and thereby provide a specific event on a screen provided on at least one display (1350). The gesture recognition cameras (1360-2, 1360-3) can recognize the user's gesture (gesture recognition), obtain a signal corresponding to the gesture, and provide a display corresponding to the signal on at least one display (1350). The processor can identify the signal corresponding to the gesture, and perform a designated function based on the identification. The gesture recognition cameras (1360-2, 1360-3) can be used to perform a spatial recognition function using SLAM and / or a depth map for 6 degrees of freedom pose (6 dof pose). The processor may perform gesture recognition and / or object tracking functions using the motion recognition cameras (1360-2, 1360-3). In one embodiment, the motion recognition cameras (1360-2, 1360-3) may be positioned on the first rim (1302-1) and / or the second rim (1302-2).

[0285] The camera (1360) included in the wearable device (1301) is not limited to the above-described gaze tracking camera (1360-1) and motion recognition cameras (1360-2, 1360-3). For example, the wearable device (1301) can identify an external object included in the user's field of view (FoV) using a camera positioned toward the FoV. The wearable device (1301) identifying an external object can be performed based on a sensor for identifying the distance between the wearable device (1301) and the external object, such as a depth sensor and / or a time of flight (ToF) sensor. The camera (1360) positioned toward the FoV can support an autofocus (AF) function and / or an optical image stabilization (OIS) function. For example, the wearable device (1301) may include a camera (1360) (e.g., a face tracking (FT) camera) positioned toward the face to obtain an image including the face of a user wearing the wearable device (1301).

[0286] Although not shown, in one embodiment, the wearable device (1301) may further include a light source (e.g., an LED) that emits light toward a subject (e.g., a user's eyes, face, and / or an external object within the FoV) being captured using the camera (1360). The light source may include an infrared wavelength LED. The light source may be disposed on at least one of the frame (1300) and the hinge units (1306, 1307).

[0287] In one embodiment, the battery module (1370) may supply power to electronic components of the wearable device (1301). In one embodiment, the battery module (1370) may be disposed within the first temple (1304) and / or the second temple (1305). For example, the battery module (1370) may be a plurality of battery modules (1370). The plurality of battery modules (1370) may be disposed within each of the first temple (1304) and the second temple (1305). In one embodiment, the battery module (1370) may be disposed at an end of the first temple (1304) and / or the second temple (1305).

[0288] The antenna module (1375) can transmit signals or power to the outside of the wearable device (1301), or receive signals or power from the outside. In one embodiment, the antenna module (1375) can be positioned within the first temple (1304) and / or the second temple (1305). For example, the antenna module (1375) can be positioned close to one surface of the first temple (1304) and / or the second temple (1305).

[0289] The speaker (1355) can output an acoustic signal to the outside of the wearable device (1301). The acoustic output module may be referred to as a speaker. In one embodiment, the speaker (1355) may be positioned within the first temple (1304) and / or the second temple (1305) so as to be positioned adjacent to the ear of a user wearing the wearable device (1301). For example, the speaker (1355) may include a second speaker (1355-2) positioned within the first temple (1304) and thus positioned adjacent to the user's left ear, and a first speaker (1355-1) positioned within the second temple (1305) and thus positioned adjacent to the user's right ear. For example, the speaker (1355) may be an example of the speaker (312) of FIG. 3A.

[0290] The light-emitting module (not shown) may include at least one light-emitting element. The light-emitting module may emit light of a color corresponding to a specific state or emit light with an action corresponding to a specific state in order to visually provide information regarding a specific state of the wearable device (1301) to the user. For example, when the wearable device (1301) requires charging, it may emit red light at a regular cycle. In one embodiment, the light-emitting module may be disposed on the first rim (1302-1) and / or the second rim (1302-2).

[0291] Referring to FIG. 13B, according to one embodiment, a wearable device (1301) may include a printed circuit board (PCB) (1390). The PCB (1390) may be included in at least one of the first temple (1304) or the second temple (1305). The PCB (1390) may include an interposer disposed between at least two sub-PCBs. One or more hardwares included in the wearable device (1301) (e.g., hardwares illustrated by different blocks in FIG. 16) may be disposed on the PCB (1390). The wearable device (1301) may include a flexible PCB (FPCB) for interconnecting the hardwares.

[0292] According to one embodiment, a wearable device (1301) may include at least one of a gyro sensor, a gravity sensor, and / or an acceleration sensor for detecting a posture of the wearable device (1301) and / or a posture of a body part (e.g., a head) of a user wearing the wearable device (1301). Each of the gravity sensor and the acceleration sensor may measure gravitational acceleration and / or acceleration based on mutually perpendicular designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). The gyro sensor may measure an angular velocity of each of the designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). At least one of the gravity sensor, the acceleration sensor, and the gyro sensor may be referred to as an inertial measurement unit (IMU). According to one embodiment, the wearable device (1301) may identify a user's motion and / or gesture performed to execute or terminate a specific function of the wearable device (1301) based on the IMU.

[0293] Referring to FIG. 13C, an embodiment of a wearable device (1301) is illustrated. As a non-limiting example, the wearable device (1301) of FIG. 13C may have the form of glasses. At least some of the hardware (or components) of the wearable device (1301) illustrated in FIGS. 13A and 13B may be applied to or included in the wearable device (1301) of FIG. 13C. Accordingly, overlapping content may be omitted.

[0294] The wearable device (1301) may include a first rim (1302-1) and / or a second rim (1302-2). The wearable device (1301) may include a display (1350). The first display (1350-1) may be disposed on the first rim (1302-1). For example, the first display (1350-1) may be disposed within the first rim (1302-1) so as to face a face of a user wearing the wearable device (1301). For example, the first display (1350-1) may be disposed within the first rim (1302-1) so as to face an eye of a user wearing the wearable device (1301). For example, the first display (1350-1) may be disposed at a location corresponding to the user's left eye. For example, the first display (1350-1) may be positioned at the upper portion within the first rim (1302-1). For example, the first display (1350-1) may display the first screen (1393-1). For example, while the wearable device (1301) is worn by the user, the user may view the first screen (1393-1). For example, the user may view the first screen (1393-1) while looking at the upper portion of the field of view.

[0295] The second display (1350-2) may be positioned on the second rim (1302-2). For example, the second display (1350-2) may be positioned within the second rim (1302-2) to face the face of a user wearing the wearable device (1301). For example, the second display (1350-2) may be positioned within the second rim (1302-2) to face the eye of a user wearing the wearable device (1301). For example, the second display (1350-2) may be positioned at a position corresponding to the user's right eye. For example, the second display (1350-2) may be positioned at an upper portion within the second rim (1302-2). For example, the second display (1350-2) may display a second screen (1393-2). For example, while the wearable device (1301) is worn by the user, the user can view the second screen (1393-2). For example, the user can view the second screen (1393-2) while looking at the upper part of the field of view.

[0296] The wearable device (1301) may include a nose pad (1310). For example, the nose pad (1310) may include a bridge (1303), a first pad (1311), and / or a second pad (1312). For example, while the wearable device (1301) is worn by a user, the first pad (1311) and the second pad (1312) may come into contact with a portion of the user's nose. For example, while the first pad (1311) and the second pad (1312) are in contact with a portion of the user's nose, the first display (1350-1) may be positioned on the first rim (1350-1) so as to face the user's face. For example, while the first pad (1311) and the second pad (1312) are in contact with a portion of the user's nose, the first display (1350-1) may be included in the user's field of view. For example, while the first pad (1311) and the second pad (1312) are in contact with a portion of the user's nose, the first screen (1393-1) may be included in the field of view of the user's left eye. For example, while the first pad (1311) and the second pad (1312) are in contact with a portion of the user's nose, the second display (1350-2) may be positioned on the second rim (1350-2) so as to face the user's face. For example, while the first pad (1311) and the second pad (1312) are in contact with a portion of the user's nose, the second display (1350-2) may be included in the field of view of the user. For example, while the first pad (1311) and the second pad (1312) are in contact with a portion of the user's nose, the second screen (1393-2) may be included in the field of view of the user's right eye.

[0297] Figures 14a and 14b illustrate an example of an exterior appearance of a wearable device. The wearable device (1301) of Figures 14a and 14b may be an example of the electronic device (301) of Figure 3a. According to one embodiment, an example of an exterior appearance of a first side (1410) of a housing of a wearable device (1301) is illustrated in Figure 14a, and an example of an exterior appearance of a second side (1420) opposite to the first side (1410) may be illustrated in Figure 14b.

[0298] Referring to FIG. 14A, a first surface (1410) of a wearable device (1301) according to one embodiment may have a form attachable to a body part of a user (e.g., the face of the user). Although not shown, the wearable device (1301) may further include a strap for fixing to a body part of a user, and / or one or more temples (e.g., the first temple (1304) and / or the second temple (1305) of FIGS. 13A to 13C). A first display (1350-1) for outputting an image to a left eye among the user's two eyes, and a second display (1350-2) for outputting an image to a right eye among the two eyes, may be disposed on the first surface (1410). The wearable device (1301) may be formed on the first surface (1410) and may further include a rubber or silicone packing to prevent interference from light (e.g., ambient light) different from the light emitted from the first display (1350-1) and the second display (1350-2).

[0299] According to one embodiment, a wearable device (1301) may include cameras (1360-1) for photographing and / or tracking both eyes of a user adjacent to each of the first display (1350-1) and the second display (1350-2). The cameras (1360-1) may be referred to as the gaze tracking camera (1360-1) of FIG. 13B. According to one embodiment, a wearable device (1301) may include cameras (1360-5, 1360-6) for photographing and / or recognizing a face of a user. The cameras (1360-5, 1360-6) may be referred to as FT cameras. The wearable device (1301) can control an avatar representing the user in a virtual space based on the facial motion of the user identified using cameras (1360-5, 1360-6). For example, the wearable device (1301) can change the texture and / or shape of a part of the avatar (e.g., a part of the avatar representing a human face) using information obtained by cameras (1360-5, 1360-6) (e.g., an FT camera) and representing the facial expression of the user wearing the wearable device (1301).

[0300] Referring to FIG. 14b, a camera (e.g., cameras (1360-7, 1360-8, 1360-9, 1360-10, 1360-11, 1360-12)) and / or a sensor (e.g., a depth sensor (1430)) for obtaining information related to the external environment of the wearable device (1301) may be disposed on a second surface (1420) opposite to the first surface (1410) of FIG. 14a. For example, the cameras (1360-7, 1360-8, 1360-9, 1360-10) may be disposed on the second surface (1420) to recognize external objects. Cameras (1360-7, 1360-8, 1360-9, 1360-10) may be referenced to the motion recognition cameras (1360-2, 1360-3) of FIG. 13b.

[0301] Using cameras (1360-11, 1360-12), the wearable device (1301) can acquire images and / or videos to be transmitted to each of the user's eyes. The camera (1360-11) can be positioned on the second face (1420) of the wearable device (1301) to acquire an image to be displayed through the second display (1350-2) corresponding to the right eye among the two eyes. The camera (1360-12) can be positioned on the second face (1420) of the wearable device (1301) to acquire an image to be displayed through the first display (1350-1) corresponding to the left eye among the two eyes. The cameras (1360-11, 1360-12) can be referred to as the shooting camera (1360-4) of FIG. 13B.

[0302] According to one embodiment, a wearable device (1301) may include a depth sensor (1430) disposed on a second face (1420) to identify a distance between the wearable device (1301) and an external object. Using the depth sensor (1430), the wearable device (1301) may obtain spatial information (e.g., a depth map) for at least a portion of a field of view (FoV) of a user wearing the wearable device (1301). Although not shown, a microphone may be disposed on the second face (1420) of the wearable device (1301) to obtain a sound output from an external object. The number of microphones may be one or more, depending on the embodiment.

[0303] Fig. 15 illustrates an example of the appearance of a wearable device. The wearable device (1301) may be an example of the electronic device (301) of Fig. 3a. For example, the wearable device (1301) illustrated in Fig. 15 may be understood as an embodiment of the wearable device (1301) illustrated in Figs. 13a to 14b.

[0304] Referring to FIG. 15, the wearable device (1301) may have a headset form and / or a headgear form. For example, the wearable device (1301) may include a speaker (1355), a microphone (1365), a strap (1530), an ear pad (1540), and / or a camera (1560). For example, the descriptions of the speaker (1355) in FIGS. 13A and 13B may be referred to for the speaker (1355). For example, the speaker (1355) may be an example of the speaker (312) in FIG. 3A. Each of the first speaker (1355-1) and the second speaker (1355-2) may be positioned to face the user's ear while the wearable device (1301) is worn by the user. For example, reference may be made to the descriptions of the microphone (1365) in FIGS. 13A and 13B for the microphone (1365). For example, the microphone (1365) may be an example of the microphone (330) in FIG. 3A. For example, the microphone (1365) may be positioned adjacent to the user's mouth while the wearable device (1301) is worn by the user. Although the microphone (1365) is illustrated in FIG. 13C as being positioned below the first camera (1560-1), the embodiment is not limited thereto. For example, the microphone (1365) may be positioned above the first camera (1560-1) and / or next to the first camera (1560-1). As a non-limiting example, the microphone (1365) may be positioned around the second camera (1560-2).

[0305] The strap (1530) may be referred to as an accessory to be attached to the user's head. For example, while the wearable device (1301) is worn by the user, the strap (1530) may come into contact with the user's head. The ear pad (1540) may be referred to as a member that comes into contact with the user's ear. The ear pad (1540) may be used to improve the wearing comfort of the wearable device (1301). The ear pad (1540) may be composed of a foam material, memory foam, fabric, and / or artificial leather. The ear pad (1540) may be used to prevent leakage of an audio signal output by the speaker (1355). The first ear pad (1540-1) may be positioned between the first speaker (1355-1) and the user's ear while the wearable device (1301) is worn by the user. The second ear pad (1540-2) can be positioned between the second speaker (1355-2) and the user's ear while the wearable device (1301) is worn by the user.

[0306] A camera (1560) can be used to acquire an image. The camera (1560) can be an example of the camera (1360) of FIGS. 13A to 14B . For example, reference may be made to the descriptions of the camera (1360) of FIGS. 13A to 14B for the camera (1560). For example, the camera (1560) can be used to acquire an image representing an external environment. For example, the wearable device (1301) can identify an external object included in an image acquired through the camera (1560). The first camera (1560-1) can be positioned to face the direction in which the user's face faces while the wearable device (1301) is worn by the user. The first camera (1560-1) can be positioned on a housing of the wearable device (1301) that includes the first speaker (1255-1). The second camera (1560-2) may be positioned to face the direction in which the user's face is facing while the wearable device (1301) is worn by the user. The second camera (1560-2) may be positioned on the housing of the wearable device (1301) including the second speaker (1255-2).

[0307] Hereinafter, with reference to FIG. 16, the hardware or software configuration of the wearable device (1301) is described.

[0308] Fig. 16 illustrates an example of a block diagram of a wearable device. The wearable device (1301) of Fig. 16 may be an example of the electronic device (301) of Fig. 3A. The wearable device (1301) may be an example of the wearable device (1301) of Figs. 13A to 15.

[0309] Referring to FIG. 16, a wearable device (1301) according to one embodiment may include a processor (1610), a memory (1615), a display (1350) (e.g., the first display (1350-1) and / or the second display (1350-2) of FIGS. 13A, 13B, 13C, 14A, and 14B), a speaker (1355), a microphone (1365), and / or at least one sensor (1620). The processor (1610), the memory (1615), the display (1350), the speaker (1355), the microphone (1365), and / or the at least one sensor (1620) may be electrically and / or operatively connected to each other by electronic components such as a communication bus (402). In the present disclosure, the operational connection of the electronic components may include a direct connection established between the electronic components and / or an indirect connection established between the electronic components such that a first electronic component among the electronic components is controlled by a second electronic component among the electronic components. The type and / or number of electronic components included in the wearable device (1301) is not limited to those illustrated in FIG. 16. For example, the wearable device (1301) may include only some of the electronic components illustrated in FIG. 16. For example, the wearable device (1301) may include the communication circuit (310) of FIG. 3A. For example, the wearable device (1301) may include the first multimodal model (350), the third model (370), and / or the fourth multimodal model (380) of FIG. 3B.

[0310] A processor (1610) of a wearable device (1301) according to one embodiment may include a circuit (e.g., a processing circuit) for processing data based on one or more instructions. The processor (1610) may be an example of at least one processor (300) of FIG. 3A. Since the processor (1610) may be substantially the same as at least one processor (300) of FIG. 3A, a redundant description thereof will be omitted. According to one embodiment, the structure of the processor (1610) is not limited to one embodiment of the present disclosure, and at least one circuit may be formed as a separate processor that is physically separated from the processor. In one embodiment including a processor (1610) having a multi-core processor architecture, operations and / or functions of the present disclosure may be individually or collectively performed by one or more cores included in the processor (1610).

[0311] According to one embodiment, the memory (1615) of the wearable device (1301) may include electronic components for storing data and / or instructions input to and / or output from the processor (1610). The memory (1615) may be an example of the memory (320) of FIG. 3A. Since the processor (1610) may be substantially the same as the memory (320) of FIG. 3A, a redundant description thereof will be omitted. In one embodiment, the memory (1615) may be referred to as storage.

[0312] In one embodiment, a display (1350) of a wearable device (1301) may output visualized information to a user of the wearable device (1301). The display (1350) may be an example of the display (311) of FIG. 3A. Since the display (1350) may be substantially the same as the display (311) of FIG. 3A, a redundant description thereof will be omitted. In addition, for the display (1350), reference may be made to the descriptions of the display (1350) of FIGS. 13A to 14B. A display (1350) arranged in front of the eyes of a user wearing a wearable device (1301) may be disposed on at least a portion of a housing of the wearable device (1301) (e.g., the first display (1350-1) and / or the second display (1350-2) of FIGS. 13A, 13B, 13C, 14A, and 14B). For example, the display (1350) may be included in a display assembly. For example, the display (1350) may be controlled by a processor (1610) to output visualized information to the user. The display (1350) may include a flexible display, a flat panel display (FPD), and / or electronic paper. The display (1350) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LED may include an OLED (organic LED). The embodiment is not limited thereto, and for example, if the wearable device (1301) includes a lens for transmitting external light (or ambient light), the display (1350) may include a projector (or projection assembly) for projecting light onto the lens.In one embodiment, the display (1350) may be referred to as a display panel and / or a display module. The pixels included in the display (1350) may be arranged to face either of the user's eyes when the wearable device (1301) is worn by the user. For example, the display (1350) may include display areas (or active areas) corresponding to each of the user's eyes.

[0313] In one embodiment, at least one sensor (1620) of the wearable device (1301) may generate electrical information that may be processed by the processor (1610) and / or the memory (1615) from non-electronic information related to the wearable device (1301). The at least one sensor (1620) may be an example of the at least one sensor (340) of FIG. 3A. The at least one sensor (1620) may be substantially the same as the at least one sensor (340) of FIG. 3A, and thus, a redundant description thereof will be omitted. For example, the at least one sensor (1620) may include a global positioning system (GPS) sensor for detecting a geographic location of the wearable device (1301). In addition to the above GPS method, at least one sensor (1620) may generate information indicating the geographic location of the wearable device (1301) based on a global navigation satellite system (GNSS) such as, for example, Galileo or Beidou (compass). The information may be stored in the memory (1615), processed by the processor (1610), and / or transmitted to another electronic device distinct from the wearable device (1301) via a communication circuit.

[0314] Referring to FIG. 16, as an example of at least one sensor (1620) included in a wearable device (1301), an image sensor (1621) and / or a motion sensor (1622) are illustrated. The image sensor (1621) may include one or more light sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate electrical signals representing the color and / or brightness of light. The image sensor (1621) may be referred to as a camera. The image sensor (1621) may include one or more cameras. For example, the image sensor (1621) may include the camera (1360) of FIGS. 13A to 14B and / or the camera (1560) of FIG. 15. For the image sensor (1621), reference may be made to the descriptions of the camera (1360) of FIGS. 13A to 14B and / or the camera (1560) of FIG. 15.

[0315] A plurality of light sensors included in the image sensor (1621) may be arranged in the form of a two-dimensional grid (2-dimensional array). The image sensor (1621) may acquire electrical signals of each of the plurality of light sensors substantially simultaneously, and generate two-dimensional frame data corresponding to light reaching the light sensors of the two-dimensional grid. For example, an image (or photograph data) captured using the image sensor (1621) may mean one (a) two-dimensional frame data acquired from the image sensor (1621). For example, video data captured using the image sensor (1621) may mean a sequence of a plurality of two-dimensional frame data acquired from the image sensor (1621) according to a frame rate. The image sensor (1621) may further include a flash light that is arranged to face a direction in which the image sensor (1621) receives light and outputs light toward the direction.

[0316] According to one embodiment, a wearable device (1301) may include a plurality of image sensors, as examples of image sensors (1621), arranged in different directions. As described above with reference to FIGS. 13A, 13B, 13C, 14A, 14B, and 15, the plurality of image sensors may include gaze tracking cameras (e.g., gaze tracking camera 1360-1 of FIGS. 13B and 14A) configured to be arranged toward the eyes of a user wearing the wearable device (1301). The plurality of image sensors may include outward cameras. The processor (1610) may identify the direction of the user's gaze using images and / or videos acquired from the gaze tracking cameras. The gaze tracking cameras may include infrared (IR) sensors. The gaze tracking cameras may be referred to as eye sensors and / or eye trackers.

[0317] The external camera may be positioned facing the front of a user wearing the wearable device (1301) (e.g., in a direction in which both eyes may face). The wearable device (1301) may include multiple external cameras. The embodiment is not limited thereto, and the external camera may be positioned facing an external space. Using images and / or videos acquired from the external cameras, the processor (1610) may identify external objects. For example, the processor (1610) may identify the position, shape, and / or gesture (e.g., hand gesture) of a hand of a user wearing the wearable device (1301) based on images and / or videos acquired from the external cameras. Using images and / or videos of the external environment acquired from the external cameras, the processor (1610) may recognize or track one or more objects within the external environment.

[0318] In one embodiment, the motion sensor (1622) may output electrical signals representing gravitational accelerations, accelerations, and / or angular velocities of a plurality of axes (e.g., x-axis, y-axis, and z-axis) that are perpendicular to each other and relative to a designated origin within the wearable device (1301) and / or the motion sensor (1622). For example, the processor (1610) may repeatedly receive or acquire sensor data from the motion sensor (1622) that includes accelerations, angular velocities, and / or magnitudes of magnetic fields of the plurality of axes based on a designated period (e.g., 1 millisecond). In one embodiment, the motion sensor (1622) may be referred to as an inertial measurement unit (IMU). At least one sensor (1620) included in the wearable device (1301) is not limited to those described above and may include a grip sensor, a proximity sensor, a heart rate sensor, a fingerprint sensor, an ambient light sensor, and / or a ToF sensor. Using the motion sensor (1622), the processor (1610) can detect motion of the wearable device (1301) (e.g., motion of the wearable device (1301) caused by a user wearing the wearable device (1301).

[0319] According to one embodiment, one or more instructions (or commands) representing data to be processed, calculations to be performed, and / or operations to be performed by the processor (1610) of the wearable device (1301) may be stored within the memory (1615) of the wearable device (1301). A set of one or more instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and / or a software application (hereinafter, “application”). For example, the wearable device (1301) and / or the processor (1610) may perform at least one of the operations of FIG. 3b, FIG. 4a, FIG. 4b, FIG. 5, FIG. 6a, FIG. 6b, FIG. 7, FIG. 8a, FIG. 8b, FIG. 9a, FIG. 9b, FIG. 10a, FIG. 10b, FIG. 11, FIG. 17, FIG. 18, FIG. 19, FIG. 20a, FIG. 20b, and FIG. 21 when a set of a plurality of instructions distributed in the form of an operating system, firmware, driver, program, and / or software application is executed. Hereinafter, the fact that a software application is installed in a wearable device (1301) may mean that one or more instructions provided in the form of a software application (or package) are stored in a memory (1615), and that the one or more applications are stored in a format executable by the processor (1610) (e.g., a file having an extension specified by the operating system of the wearable device (1301)). As an example, the application may include a program and / or a library related to a service provided to a user.

[0320] Referring to FIG. 16, programs installed in the wearable device (1301) may be included in any one of different layers, including an application layer (1640), a framework layer (1650), and / or a hardware abstraction layer (HAL) (1680), based on the target. For example, programs (e.g., modules or drivers) designed to target the hardware (e.g., the display (1350), and / or at least one sensor (1620)) of the wearable device (1301) may be included in the hardware abstraction layer (1680) (e.g., the android system HAL, and / or the XR HAL). The framework layer (1650) may be referred to as an XR framework layer in the sense that it includes one or more programs for providing an XR (extended reality) service. For example, the layers illustrated in FIG. 16 may be logically (or for convenience of explanation) separated, and may not mean that the address space of the memory (1615) is separated by the layers.

[0321] Within the framework layer (1650), programs designed to target at least one of the hardware abstraction layer (1680) and / or the application layer (1640) (e.g., a position tracker (1671), a space recognizer (1672), a gesture tracker (1673), an eye tracker (1674), a face tracker (1675), and / or a renderer (1690)) may be included. The programs included in the framework layer (1650) may provide an API (application programming interface) that is executable (or callable) based on other programs.

[0322] The application layer (1640) may include programs designed to target users of the wearable device (1301). Examples of programs included in the application layer (1640) include, but are not limited to, an extended reality (XR) system user interface (UI) (1641) and / or an XR application (1642). For example, programs (e.g., software applications) included in the application layer (1640) may call APIs to cause execution of functions supported by programs included in the framework layer (1650).

[0323] The wearable device (1301) may display one or more visual objects on the display (1350) for performing interaction with the user based on the execution of the XR system UI (1641). A visual object may refer to an object that can be placed on a screen for transmitting and / or interacting with information, such as text, an image, an icon, a video, a button, a checkbox, a radio button, a text box, a slider, and / or a table. A visual object may be referred to as a visual guide, a virtual object, a visual element, a UI element, a view object, and / or a view element. The wearable device (1301) may provide the user with functions available within a virtual space based on the execution of the XR system UI (1641).

[0324] Referring to FIG. 16, a lightweight renderer (1643) and / or an XR plug-in (1644) are illustrated to be included within the XR system UI (1641), but are not limited thereto. For example, based on the XR system UI (1641), the processor (1610) may execute a lightweight renderer (1643) and / or an XR plug-in (1644) within the framework layer (1650).

[0325] The wearable device (1301) may acquire resources (e.g., APIs, system processes, and / or libraries) used to define, create, and / or execute a rendering pipeline that allows partial changes based on the execution of a lightweight renderer (1643). The lightweight renderer (1643) may be referred to as a lightweight render pipeline in terms of defining a rendering pipeline that allows partial changes. The lightweight renderer (1643) may include a renderer built prior to the execution of a software application (e.g., a prebuilt renderer). For example, the wearable device (1301) may acquire resources (e.g., APIs, system processes, and / or libraries) used to define, create, and / or execute an entire rendering pipeline based on the execution of an XR plug-in (1644). The XR plugin (1644) can be referred to as an open XR native client from the perspective of defining (or configuring) the entire rendering pipeline.

[0326] The wearable device (1301) may display a screen representing at least a portion of a virtual space on the display (1350) based on the execution of the XR application (1642). The XR plug-in (1644-1) included in the XR application (1642) may include instructions that support functions similar to those of the XR plug-in (1644) of the XR system UI (1641). Descriptions of the XR plug-in (1644-1) that overlap with those of the XR plug-in (1644) may be omitted. The wearable device (1301) may cause the execution of the virtual space manager (1651) based on the execution of the XR application (1642).

[0327] The wearable device (1301) can display an image on the display (1350) in a virtual space based on the execution of the application (16165). The application (16165) can be configured to output image information for displaying a two-dimensional image. The wearable device (1301) can cause the execution of the virtual space manager (1651) based on the execution of the application (16165). The wearable device (1301) can generate dual image information to display the two-dimensional image in a three-dimensional virtual space based on the execution of the application (16165). Here, the dual image information can include first image information for the left eye and second image information for the right eye, taking into account binocular disparity. In order to display the two-dimensional image in the three-dimensional virtual space, the wearable device (1301) can generate the dual image information based on the image information for displaying the two-dimensional image.

[0328] According to one embodiment, the wearable device (1301) can provide a virtual space service based on the execution of the virtual space manager (1651). For example, the virtual space manager (1651) can include a platform for supporting the virtual space service. Based on the execution of the virtual space manager (1651), the wearable device (1301) can identify a virtual space formed based on the user's location indicated by data acquired through at least one sensor (1620), and can display at least a portion of the virtual space on the display (1350). The virtual space manager (1651) can be referred to as a composition presentation manager (CPM).

[0329] The virtual space manager (1651) may include a runtime service (1652). For example, the runtime service (1652) may be referred to as an OpenXR runtime module (or an OpenXR runtime program). The wearable device (1301) may execute at least one of a user's pose prediction function, a frame timing function, and / or a spatial input function based on the execution of the runtime service (1652). For example, the wearable device (1301) may perform rendering for a virtual space service for the user based on the execution of the runtime service (1652). For example, a function related to a virtual space, executable by the application layer (1640), may be supported based on the execution of the runtime service (1652).

[0330] The virtual space manager (1651) may include a pass-through manager (1653). Based on the execution of the pass-through manager (1653), the wearable device (1301) may display an image and / or video representing an actual space acquired through an external camera on at least a portion of the screen while displaying a screen representing a virtual space on the display (1350).

[0331] The virtual space manager (1651) may include an input manager (1654). The wearable device (1301) may identify data (e.g., sensor data) acquired by executing one or more programs included in the recognition service layer (1670) based on the execution of the input manager (1654). The wearable device (1301) may use the acquired data to identify user input related to the wearable device (1301). The user input may be related to a motion (e.g., a hand gesture), gaze, and / or speech of the user identified by at least one sensor (1620) (e.g., an image sensor (1621) such as an external camera). The user input may be identified based on an external electronic device connected (or paired) via a communication circuit.

[0332] The perception abstract layer (1660) can be used for data exchange between the virtual space manager (1651) and the perception service layer (1670). From the perspective of being used for data exchange between the virtual space manager (1651) and the perception service layer (1670), the perception abstract layer (1660) can be referred to as an interface. For example, the perception abstract layer (1660) can be referenced as OpenPX. The perception abstract layer (1660) can be used for a perception client and a perception service.

[0333] According to one embodiment, the recognition service layer (1670) may include one or more programs for processing data acquired from at least one sensor (1620). The one or more programs may include at least one of a position tracker (1671), a spatial recognizer (1672), a gesture tracker (1673), an eye tracker (1674), a face tracker (1675), and / or a renderer (1690). The type and / or number of the one or more programs included in the recognition service layer (1670) are not limited to those illustrated in FIG. 16.

[0334] The wearable device (1301) can identify the pose of the wearable device (1301) using at least one sensor (1620) based on the execution of the position tracker (1671). The wearable device (1301) can identify the 6 degrees of freedom pose (6 dof pose) of the wearable device (1301) using data acquired using an external camera (e.g., an image sensor (1621)) and / or an IMU (e.g., a motion sensor (1622) including a gyro sensor, an acceleration sensor, and / or a geomagnetic sensor) based on the execution of the position tracker (1671). The position tracker (1671) can be referred to as a head tracking (HeT) module (or head tracker, a head tracking program).

[0335] The wearable device (1301) can obtain information for providing a three-dimensional virtual space corresponding to the surrounding environment (e.g., external space) of the wearable device (1301) (or the user of the wearable device (1301)) based on the execution of the space recognizer (1672). The wearable device (1301) can reproduce the surrounding environment of the wearable device (1301) in three dimensions using data obtained using an external camera (e.g., an image sensor (1621)) based on the execution of the space recognizer (1672). The wearable device (1301) can identify at least one of a plane, a slope, and stairs based on the surrounding environment of the wearable device (1301) reproduced in three dimensions based on the execution of the space recognizer (1672). The spatial recognizer (1672) may be referred to as a scene understanding (SU) module (or scene recognition program).

[0336] The wearable device (1301) can identify (or recognize) a pose and / or gesture of a hand of a user of the wearable device (1301) based on the execution of the gesture tracker (1673). For example, the wearable device (1301) can identify a pose and / or gesture of a hand of a user using data acquired from an external camera (e.g., an image sensor (1621)) based on the execution of the gesture tracker (1673). For example, the wearable device (1301) can identify a pose and / or gesture of a hand of a user based on data (or images) acquired using an external camera based on the execution of the gesture tracker (1673). The gesture tracker (1673) may be referred to as a hand tracking (HaT) module (or hand tracking program) and / or a gesture tracking module.

[0337] The wearable device (1301) can identify (or track) eye movements of a user of the wearable device (1301) based on the execution of the gaze tracker (1674). For example, the wearable device (1301) can identify eye movements of the user using data acquired from a gaze tracking camera (e.g., an image sensor (1621)) based on the execution of the gaze tracker (1674). The gaze tracker (1674) may be referred to as an eye tracking (ET) module (or eye tracking program) and / or a gaze tracking module.

[0338] The recognition service layer (1670) of the wearable device (1301) may further include a face tracker (1675) for tracking the user's face. For example, the wearable device (1301) may identify (or track) the movement of the user's face and / or the user's expression based on the execution of the face tracker (1675). The wearable device (1301) may estimate the user's expression based on the movement of the user's face based on the execution of the face tracker (1675). As an example, the wearable device (1301) may identify the movement of the user's face and / or the user's expression based on data (e.g., images and / or videos) acquired using an FT camera (e.g., a camera facing at least a portion of the user's face, an image sensor (1621)) based on the execution of the face tracker (1675). The face tracker (1675) may be referred to as a face tracking (FT) (or face tracking program), and / or a face tracking module.

[0339] The renderer (1690) may include instructions for rendering images in a three-dimensional virtual space. The processor (1610) (e.g., DPU) executing the renderer (1690) may obtain at least one image to be at least partially displayed in the display area of ​​the display (1350) from a software application (e.g., a software application executed by the CPU and / or GPU). For example, the processor (1610) executing the renderer (1690) may determine the location of the area in which the application (e.g., XR application (1642), application (16165)) is to be rendered. The processor (1610) executing the renderer (1690) may generate an image of the application to be displayed on the display (1350). The renderer (1690) may synthesize images to generate a composite image to be displayed on the display (1350).

[0340] The processor (1610) executing the renderer (1690) can divide the display area of ​​the display (1350) into a foveated portion (or may be referred to as a foveated area) and a peripheral portion (or may be referred to as a residual area) using the gaze position calculated using the position tracker (1671) and / or the gaze tracker (1674). For example, the processor (1610) detecting the coordinate values ​​of the gaze position can determine the portion of the display area including the coordinate values ​​as the foveated area. The DPU executing the renderer (1690) can obtain at least one image corresponding to each of the foveated area and the residual area, and having a size smaller than the size of the entire display area of ​​the display (1350) or a resolution smaller than the resolution of the display area.

[0341] The processor (1610) executing the renderer (1690) may obtain or generate a composite image to be displayed on the display (1350) by synthesizing an image corresponding to the foveated area and an image corresponding to the surrounding area. For example, the processor (1610) may perform upscaling to enlarge the image corresponding to the surrounding area to the size of the entire display area of ​​the display (1350). On the enlarged image, the processor (1610) may combine the image corresponding to the foveated area to generate a composite image to be displayed on the display (1350). Along the boundary line of the image corresponding to the foveated area, the processor (1610) may apply a visual effect, such as blur, to blend the enlarged image and the image corresponding to the foveated area.

[0342] Fig. 17 shows an example of a block diagram of a wearable device for displaying an image in a virtual space. The wearable device of Fig. 17 (e.g., wearable device (1301)) may be an example of the electronic device (301) of Fig. 3A. Fig. 17 describes an example of executing multiple programs (or instructions) for displaying an image in a virtual space. The multiple programs (or instructions) may all be executed in one processor (e.g., AP) or may be executed by multiple processors (e.g., AP, GPU (graphics processing unit), NPU (neural processing unit)). The meaning of being executed by the multiple processors means that some programs (or instructions) may be executed by a first processor and other programs (or instructions) may be executed by a second processor different from the first processor.

[0343] Referring to FIG. 17, the wearable device (1301) may execute a virtual space manager (1750) (e.g., the virtual space manager (1651) of FIG. 16, CPM) to render an image in a virtual space. For the virtual space manager (1750), at least some of the descriptions of the virtual space manager (1651) of FIG. 16 may be referenced. The virtual space manager (1750) may include a platform for supporting a virtual space service. The virtual space manager (1750) may include a runtime service (1751) (e.g., OpenXR Runtime), a panel rendering (1752) (e.g., 2D Panel Render), and an XR compositor (1753). The wearable device (1301) may execute at least one of a user's pose prediction function, a frame timing function, and / or a spatial input function based on the execution of the runtime service (1751). For the runtime service (1751), at least some of the descriptions of the runtime service (1652) of FIG. 16 may be referred to. The wearable device (1301) may display at least one image (video) on a panel (e.g., a 2D panel) to implement a virtual space through the display (1350) based on the execution of the panel rendering (1752). For example, the wearable device (1301) may display a rendering image corresponding to RGB information (1766) for the panel from the spatialization manager (1740) described below through the display (e.g., the display (1350)). The wearable device (1301) may synthesize an image of an actual area captured by a camera in the virtual space (hereinafter, a pass-through image) with an image of a virtual area based on the execution of the XR synthesis unit (1753) (XR Compositor). For example, the wearable device (1301) can generate a composite image by merging the pass-through image and the virtual area image based on the execution of the XR synthesis unit (1753).The wearable device (1301) can transmit the generated composite image to the display buffer so that the composite image is displayed. The wearable device (1301) can identify a virtual space through a virtual space manager (1750) and display at least a portion of the virtual space on the display (1350). The virtual space manager (1750) may be referred to as a CPM. The wearable device (1301) can execute the virtual space manager (1750) to render an image corresponding to at least a portion of the virtual space.

[0344] In one embodiment, the wearable device (1301) may execute a spatialization manager (1740). The spatialization manager (1740) may perform processes for displaying an image in a three-dimensional virtual space. The wearable device (1301) may perform preprocessing based on the execution of the spatialization manager (1740) so that the image can be rendered in a three-dimensional virtual space through the virtual space manager (1750). For example, the wearable device (1301) may perform at least some of the functions of the renderer (1690) of FIG. 16 based on the execution of the spatialization manager (1740). The wearable device (1301) can process image information provided by an application (e.g., an XR application (1710), an application providing a general 2D screen other than XR (1720), an application providing a system UI (1730)) based on the execution of the spatialization manager (1740). The spatialization manager (1740) (e.g., Space Flinger) can include a system scene manager (1741) (e.g., System scene), an input manager (1742) (e.g., Input Routing), and a lightweight rendering engine (1743) (e.g., Impress Engine). The system scene manager (1741) can be executed to display the system UI (1730). System UI-related information (1764) can be transmitted to the system scene manager (1741) from a program (e.g., API) providing the system UI (1730). System UI-related information (1764) can be obtained via a spatializer API and / or a same-process private API. The spatialization manager (1740) can determine the screen layout (e.g., location, display order) of the system UI (1730) in three-dimensional space using pre-allocated resources.The system screen manager (1741) may transmit image information (1767) for rendering the screen of the system UI (1730) to the virtual space manager (1750) according to the layout. The input manager (1742) may be configured to process user input (e.g., user input on a system screen or an app screen). The input manager (1742) may map user input recognized by at least one sensor (1620) of the wearable device (1301) to at least one of one or more software applications mapped to a virtual space by the spatialization manager (1740) (e.g., an XR application (1710), an application providing a general 2D screen other than XR (1720), an application providing a system UI (1730)). For example, the mapping of the user input may include an operation of executing instructions (e.g., a subroutine and / or an event handler) of a software application for processing the user input. The lightweight rendering engine (1743) may be a renderer for generating images (e.g., a lightweight renderer (1643)). For example, the lightweight rendering engine (1743) may be used to display a system UI (1730).

[0345] In one embodiment, the spatialization manager (1740) may include a lightweight rendering engine (1743) for rendering the system UI. In one embodiment, if the lightweight rendering engine (1743) does not have sufficient resources to render an avatar used in the HMD, at least one external rendering engine may be used. In this case, to resolve compatibility issues with external rendering (e.g., a 3rd party engine), an external rendering engine support module may be added within the spatialization manager (1740).

[0346] According to one embodiment, the electronic device can execute an application. For example, in response to the execution of an XR application (1710) (e.g., an XR application (1642), a 3D game, an XR map, or other immersive application), the electronic device can execute a virtual space manager (1750). The wearable device (1301) can provide dual image information (1761) provided from the XR application (1710) to the virtual space manager (1750). In order to display an image in a three-dimensional space, the dual image information (1761) can include two pieces of image information that take binocular parallax into account. For example, the dual image information (1761) can include first image information for the user's left eye and second image information for the user's right eye for rendering in a three-dimensional virtual space. Hereinafter, in the present disclosure, the term dual image information is used to refer to image information for displaying images for both eyes in a three-dimensional space. In addition to dual image information, the above dual image information may also include binocular image information, dual image data, dual images, binocular image data, stereoscopic image information, 3D image information, spatial image information, spatial image data, 2D-3D conversion data, dimensional conversion image data, binocular parallax image data, and / or equivalent technical terms. The wearable device (1301) can generate a composite image by merging image layers through a virtual space manager (1750). The wearable device (1301) can transmit the generated composite image to a display buffer. The composite image can be displayed on the display (1350) of the wearable device (1301).

[0347] According to one embodiment, the electronic device can execute at least one of an XR application (1710) and other applications (1720) (e.g., a first application (1720-1), a second application (1720-2), ..., an Nth application (1720-N)). According to one embodiment, the application (1720) can be configured to output image information for displaying a two-dimensional (2D) image. In other words, the application (1720) can provide a two-dimensional (2D) image (e.g., a window and / or an activity). As an example, the application (1720) can be a video application, a schedule application, or an Internet browser application. If, in response to the execution of the application (1720), the image information (1762) provided from the application (1720) is provided to the virtual space manager (1750), the image information (1762) only has x-coordinates and y-coordinates within a two-dimensional plane, so it may be difficult to consider the chronological relationship (i.e., distance from the user) between other applications centered on the user. The wearable device (1301) may execute the spatialization manager (1740) to provide dual image information to the virtual space manager (1750) even when displaying the application (1720) that provides a general 2D screen. For example, based on the execution of the spatialization manager (1740), the wearable device (1301) may receive application-related information (1763) from the first application (1720-1). For example, application-related information (1763) may include image information representing a two-dimensional image of the first application (1720-1) (e.g., information including RGB per pixel) and / or content information in the first application (1720-1) (e.g., characteristics of content executed in the first application, type of content). The application-related information (1763) may be obtained through a spatializer API.Based on the execution of the spatialization manager (1740), the wearable device (1301) can identify information about the location of the area to be rendered by the first application (1720-1) and the size of the area to be rendered (hereinafter, location information). Based on the execution of the spatialization manager (1740), the wearable device (1301) can generate dual image information (1765, e.g., RGBx2) that takes into account the user's binocular disparity through the image information and the location information. Based on the execution of the spatialization manager (1740), the wearable device (1301) can provide the dual image information (1765) to the virtual space manager (1750). By converting a simple two-dimensional image into the dual image information (1765), a problem that occurs when the image information (1762) is directly transmitted to the virtual space manager (1750) can be resolved. Additionally, since at least some of the functions for displaying images in a virtual space are performed by the spatialization manager (1740) instead of the virtual space manager (1750), the burden on the virtual space manager (1750) can be reduced.

[0348] Fig. 18 illustrates an example of a wearable device that recognizes audio signals emitted from an external object. The operations of the wearable device (1301) illustrated in Fig. 18 may be executed, performed, or controlled by a processor (e.g., processor (1610)).

[0349] Referring to FIG. 18, a user (1801) may wear a wearable device (1301). The wearable device (1301) may acquire an audio signal through a microphone (e.g., microphone (1365)). For example, the audio signal may include a voice signal (1810) and / or noise. For example, the size of the voice signal (1810) may be relatively small within the audio signal. For example, the size of the noise within the audio signal may be relatively large. For example, the SNR of the audio signal may be relatively small. For example, when the SNR of the audio signal is relatively small, a method for increasing the recognition rate of the voice signal (1810) may be required.

[0350] In one embodiment, the wearable device (1301) can detect a change in the face of the speaker (1802) of the voice signal (1810). For example, the change in the face of the speaker (1802) can be referred to as a change in the shape of the face caused by the speaker (1802) speaking. For example, the speaker (1802) can be referred to as an utterer. For example, the wearable device (1301) can detect a change in the lips of the speaker (1802) of the voice signal (1810). For example, the change in the lips of the speaker (1802) can be referred to as a change in the shape of the lips caused by the speaker (1802) speaking. For example, the wearable device (1301) can obtain sensing data through at least one sensor (e.g., at least one sensor (1620)). For example, the sensing data can include an image obtained through an image sensor (e.g., the image sensor (1621)). For example, the image can be an image of an external environment (or a physical environment) captured. For example, the image can include a visual object (1821) corresponding to a speaker (1802). For example, a part of the visual object (1821) can correspond to the face (e.g., lips) of the speaker (1802). For example, the wearable device (1301) can detect a change in the face (e.g., lips) of the speaker (1802) by identifying a part of the visual object (1821) in the image obtained through the image sensor (1621).

[0351] The wearable device (1301) can provide an audio signal (e.g., including a voice signal (1810)) acquired through a microphone (1365) and / or sensing data through at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can provide an audio signal and / or data indicating changes in the lips of a speaker (1802) to the first multimodal model (350). For example, the wearable device (1301) can identify a voice signal (1810) included in an audio signal even in an environment where the SNR of the audio signal is relatively low by using the first multimodal model (350). For example, the wearable device (1301) can identify the voice signal (1810) of a speaker (1802) even if the voice signal is relatively noisy and / or the speaker (1802) is located relatively far away.

[0352] The wearable device (1301) can obtain response information for the audio signal and / or sensing data by providing the audio signal and / or sensing data to the first multimodal model (350). For example, the response information can be generated by the first multimodal model (350) according to the audio signal and / or sensing data. For example, the response information can be used to display text representing a voice signal (1810) included in the audio signal through a display (e.g., the display (1350)). For example, the response information can be used to output an audio signal representing the voice signal (1810) through a speaker (e.g., the speaker (1355)). As a non-limiting example, the audio signal provided to the first multimodal model (350) can be an audio signal on which noise cancellation is performed according to a fourth multimodal model (e.g., the fourth multimodal model (380)).

[0353] The wearable device (1301) can provide a pass-through function for the external environment by using an image acquired through the image sensor (1621). For example, the wearable device (1301) can provide a screen (1820) representing at least a portion of a virtual space (or a three-dimensional space) corresponding to a physical environment through a display (e.g., a display (1350)). For example, the screen (1820) can include a visual object (1821). For example, the visual object (1821) can correspond to a speaker (1802). For example, the visual object (1821) can represent the speaker (1802).

[0354] The wearable device (1301) may use the response information to display text (1827) in an area (1825) within the screen (1820). The text (1827) may correspond to the voice signal (1810). For example, the text (1827) may represent the voice signal (1810). The location of the area (1825) may be adjacent to the visual object (1821). For example, the area (1825) may be parallel to the visual object (1821) on the screen (1820). For example, the area (1825) may have a predefined location on the screen (1820). As a non-limiting example, the area (1825) may overlap at least a portion of the visual object (1821).

[0355] In one embodiment, the wearable device (1301) may use the response information to output an audio signal (1830) through the speaker (1355). For example, the audio signal (1830) may correspond to the voice signal (1810). For example, the SNR of the audio signal (1830) may be higher than the SNR of an audio signal (e.g., including the voice signal (1810)) acquired through the microphone (1365). In one embodiment, the audio signal (1830) may be acquired by performing noise canceling on the audio signal acquired through the microphone (1365). For example, the fourth multimodal model (380) may be used for noise canceling. In one embodiment, the audio signal (1830) may be generated by the electronic device (1301). For example, an audio signal (1830) may be generated using the response information. For example, the wearable device (1301) may generate or obtain the audio signal (1830) by performing text-to-speech (TTS) on text (1827) representing the voice signal (1810). For example, the audio signal (1830) may be generated by the first multimodal model (350). For example, the voice feature of the generated audio signal (1830) may be a predetermined voice feature. For example, the voice feature of the generated audio signal (1830) may be substantially the same as an audio signal of an audio signal obtained through a microphone (1365) (e.g., including the voice signal (1810)). For example, the electronic device (1301) may generate the audio signal (1830) using the voice feature of the audio signal obtained through the microphone (1365). As a non-limiting example, the electronic device (1301) may generate an audio signal (1830) using pre-stored voice characteristics of another person (e.g., a celebrity).For example, the voice features may include at least one of the gender, glottal morphology, timbre, pitch, or pronunciation of the speaker (1802).

[0356] As a non-limiting example, there may be a region of a voice signal (1810) that is not identified by the wearable device (1301) in the audio signal acquired through the microphone (1365). For example, the wearable device (1301) may perform inference on the region. For example, the wearable device (1301) may perform inference on the region using an artificial intelligence model. For example, the wearable device (1301) may obtain text (1827) and / or an audio signal (1830) by performing inference on the region.

[0357] In one embodiment, the wearable device (1301) may display a visual object (1829) through the screen (1820). For example, the visual object (1829) may indicate that the wearable device (1301) provides a function to display text (1827) through the display (1350) based on an audio signal (e.g., a voice signal (1810)) acquired through the microphone (1365). Additionally, the visual object (1829) may indicate that the wearable device (1301) provides a function to output an audio signal (1830) through the speaker (1355) based on an audio signal (e.g., a voice signal (1810)) acquired through the microphone (1365). For example, the wearable device (1301) may change the state of the wearable device (1301) to a state that does not provide a function of displaying text (1827) through the display (1350) based on a user input. For example, the wearable device (1301) may change the state of the wearable device (1301) to a state that does not provide a function of outputting an audio signal (1830) through the speaker (1355) based on a user input. For example, the user input may include an input to a visual object (1829), but the embodiment is not limited thereto. For example, the user input may include a predetermined gesture input.

[0358] According to one embodiment, the wearable device (1301) may provide a function of displaying text (1827) based on an audio signal through a display (1350) and / or a function of outputting an audio signal (1830) based on the audio signal through a speaker (1355) while the wearable device (1301) is positioned in an environment where noise is greater than a reference noise. The wearable device (1301) may provide a function of displaying text (1827) based on an audio signal through a display (1350) and / or a function of outputting an audio signal (1830) based on the audio signal through a speaker (1355) when the quality of the audio signal acquired through the microphone (1365) is relatively low. For example, the wearable device (1301) may identify the size of a voice signal (1810) included in an audio signal acquired through the microphone (1365). For example, the wearable device (1301) may provide a function to display text (1827) through a display (1350) based on an audio signal, based on identifying that the size of the voice signal (1810) is smaller than a reference size. For example, the wearable device (1301) may provide a function to output an audio signal (1830) through a speaker (1355) based on an audio signal, based on identifying that the size of the voice signal (1810) is smaller than a reference size.

[0359] The wearable device (1301) may refrain from, stop, or bypass providing a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the size of the voice signal (1810) is greater than a reference size. For example, the wearable device (1301) may refrain from, stop, or bypass providing a function to output an audio signal (1830) through the speaker (1355) based on the audio signal based on identifying that the size of the voice signal (1810) is greater than a reference size.

[0360] According to one embodiment, the wearable device (1301) can identify the SNR of an audio signal (e.g., including a voice signal (1810)) acquired via the microphone (1365). For example, the wearable device (1301) can provide a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the SNR of the audio signal is less than a reference value. For example, the wearable device (1301) can provide a function to output the audio signal (1830) through the speaker (1355) based on the audio signal based on identifying that the SNR is less than a reference value.

[0361] The wearable device (1301) may refrain from, stop, or bypass providing a function to display text (1827) through the display (1350) based on the audio signal based on identifying that the SNR of the audio signal of the voice signal (1810) is greater than a reference value. For example, the wearable device (1301) may refrain from, stop, or bypass providing a function to output the audio signal (1830) through the speaker (1355) based on the audio signal based on identifying that the SNR of the audio signal is greater than a reference value.

[0362] According to one embodiment, the wearable device (1301) can communicate with an external electronic device (1838) (e.g., a smartphone). For example, the external electronic device (1838) may be an example of the electronic device (301) of FIG. 3A and / or FIG. 3B . For example, the wearable device (1301) may be referred to as a companion device of the external electronic device (1838). For example, the wearable device (1301) may establish a communication link between the wearable device (1301) and the external electronic device (1838). For example, the wearable device (1301) may transmit an audio signal (e.g., including a voice signal (1810)) acquired via a microphone (1365) to the external electronic device (1838). For example, the wearable device (1301) may transmit sensing data acquired through at least one sensor (1620) to an external electronic device (1838). For example, the sensing data may include an image acquired through an image sensor (1621). For example, the image may include a portion of a visual object (1821) corresponding to the lips of a speaker (1802).

[0363] According to one embodiment, the external electronic device (1838) may include the first multimodal model (350). For example, the external electronic device (1838) may display a screen (1840) through the display of the external electronic device (1838) based on an audio signal transmitted from the wearable device (1301) and / or sensing data transmitted from the wearable device (1301). For example, the screen (1840) may correspond to the screen (1820). For example, the screen (1840) may be substantially identical to the screen (1820), and thus, a redundant description will be omitted. For example, the visual object (1841) may correspond to the visual object (1821). For example, the area (1845) may correspond to the area (1825). For example, the descriptions for the area (1825) may be referenced for the area (1845). For example, text (1847) may correspond to text (1827). For text (1827), reference may be made to descriptions of text (1847). The external electronic device (1838) may output an audio signal (1850) through a speaker of the external electronic device (1838) based on the audio signal and / or the transmitted sensing data. For example, the audio signal (1850) may correspond to audio signal (1830). For audio signal (1830), reference may be made to descriptions of audio signal (1850).

[0364] In one embodiment, the external electronic device (1838) can transmit data representing text (1847) and / or data representing audio signals (1850) to the wearable device (1301). For example, the wearable device (1301) can display text (1827) through the display (1350) using the data representing text (1847). For example, the wearable device (1301) can output audio signals (1830) through the speaker (1355) using the data representing audio signals (1850).

[0365] In FIG. 18, the wearable device (1301) is illustrated as providing a pass-through function, but the embodiment is not limited thereto. For example, the wearable device (1301) may perform the operations illustrated in FIG. 18 while displaying or playing a video through the display (1350).

[0366] FIG. 19 illustrates an example of a wearable device that adaptively provides a response to an audio signal based on the gaze of an external object. The operations of the wearable device (1301) illustrated in FIG. 19 may be executed, performed, or controlled by a processor (e.g., processor (1610)). In the description of FIG. 19, any content that overlaps with the description of FIG. 18 may be omitted.

[0367] Referring to FIG. 19, in example (1901), a wearable device (1301) can provide a pass-through function to a user (1801). The wearable device (1301) can obtain an audio signal through a microphone (e.g., microphone (1365)). For example, the audio signal can include a voice signal (1810), a voice signal (1905), and / or noise. The wearable device (1301) can obtain sensing data through at least one sensor (e.g., at least one sensor (1620)). For example, the sensing data can include an image obtained through an image sensor (e.g., image sensor (1621)). For example, the image can be an image of an external environment (or physical environment) captured. For example, the image may include a visual object (1921) corresponding to the speaker (1802) and / or a visual object (1923) corresponding to the speaker (1903). For example, the visual object (1921) may be an example of the visual object (1821) of FIG. 18. For example, the image may include a portion of the visual object (1921) corresponding to the lips of the speaker (1802) and / or a portion of the visual object (1923) corresponding to the lips of the speaker (1903).

[0368] The wearable device (1301) can provide audio signals acquired via a microphone (1365) and / or sensing data via at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can identify a voice signal (1810) and / or a voice signal (1905) using the first multimodal model (350) even in an environment where the SNR of the audio signal is relatively low.

[0369] The wearable device (1301) can obtain response information for the audio signal and / or the sensing data by providing the audio signal and / or the sensing data to the first multimodal model (350). For example, the response information can be generated by the first multimodal model (350) according to the audio signal and / or the sensing data. For example, the response information can be used to display text representing the voice signal (1810) and / or text representing the voice signal (1905) through a display (e.g., the display (1350)). For example, the response information can be used to output an audio signal representing the voice signal (1810) and / or an audio signal representing the voice signal (1905) through a speaker (e.g., the speaker (1355)). As a non-limiting example, the audio signal provided to the first multimodal model (350) may be an audio signal on which noise cancellation has been performed according to the fourth multimodal model (e.g., the fourth multimodal model (380)).

[0370] In one embodiment, the wearable device (1301) may display a screen (1920) via the display (1350). The screen (1920) may include a visual object (1921) corresponding to a speaker (1802) and / or a visual object (1923) corresponding to a speaker (1903). For example, the wearable device (1301) may use the response information to display text (1927) in an area (1925) within the screen (1920). For example, the text (1927) may represent a voice signal (1810). For example, the wearable device (1301) may use the response information to display text (1928) in an area (1926) within the screen (1920). For example, the text (1928) may represent a voice signal (1905). For example, each of area (1925) and area (1926) may be substantially identical to area (1825) of FIG. 18, so overlapping content may be omitted.

[0371] The wearable device (1301) may use the response information to output an audio signal (1930) and / or an audio signal (1931) through a speaker (e.g., speaker (1355)). For example, the audio signal (1930) may represent the voice signal (1810). For example, the quality of the audio signal (1930) may be higher than the quality of the voice signal (1810). For example, the SNR of the audio signal (1930) may be higher than the SNR of the voice signal (1810). For example, the size (e.g., volume) of the audio signal (1930) may be greater than the size of the voice signal (1810). For example, the audio signal (1931) may represent the voice signal (1905). For example, the quality of the audio signal (1931) may be higher than the quality of the voice signal (1905). For example, the SNR of the audio signal (1931) may be higher than the SNR of the voice signal (1905). For example, the magnitude (e.g., loudness) of the audio signal (1931) may be greater than the magnitude of the voice signal (1905).

[0372] In one embodiment, the wearable device (1301) can identify the location of the speaker (1802) and / or the location of the speaker (1903) using sensing data acquired through at least one sensor (1620). For example, the sensing data may include an image acquired through an image sensor (1621). For example, the image may include a visual object (1921) and / or a visual object (1923). As a non-limiting example, the sensing data may include data acquired through a ToF sensor. For example, the wearable device (1301) can identify the distance between the wearable device (1301) and the speaker (1802). For example, the wearable device (1301) can identify the direction from the wearable device (1301) toward the speaker (1802). The wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1802) is facing. For example, the wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1802) is facing by identifying the direction in which a portion corresponding to the face (e.g., lips) of a visual object (1921) is facing.

[0373] The wearable device (1301) can identify the distance between the wearable device (1301) and the speaker (1903). For example, the wearable device (1301) can identify the direction from the wearable device (1301) toward the speaker (1903). The wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1903) is facing. For example, the wearable device (1301) can identify the direction in which the face (e.g., lips) (or gaze) of the speaker (1903) is facing by identifying the direction in which a portion corresponding to the face (e.g., lips) of a visual object (1923) is facing.

[0374] The wearable device (1301) can identify a voice signal (e.g., voice signal (1810), voice signal (1905)) included in the audio signal by using an audio signal acquired through a microphone (1365). For example, the wearable device (1301) can identify the speech location of the voice signal included in the audio signal. For example, the microphone (1365) can include a plurality of microphones. For example, each of the plurality of microphones can acquire a voice signal. For example, the phase of the acquired voice signal can be different depending on the time difference of the voice signal reaching each of the plurality of microphones. For example, the wearable device (1301) can identify the speech location of the voice signal by using the phase of the acquired voice signal. As a non-limiting example, the wearable device (1301) can identify the location of speech of a voice signal by utilizing the difference in sound pressure of the voice signal reaching each of a plurality of microphones.

[0375] The wearable device (1301) can identify the speaker of each of the voice signals (e.g., voice signal (1810), voice signal (1905)). The wearable device (1301) can determine the speaker of the voice signal based on the similarity between the location of the identified voice signal and the location of the identified speaker (e.g., speaker (1802), speaker (1903)). For example, the wearable device (1301) can identify that the voice signal (1810) is spoken by the speaker (1802) based on the similarity between the location of the voice signal (1810) and the location of the speaker (1802). For example, the wearable device (1301) can identify that the voice signal (1905) is uttered by the speaker (1903) based on the similarity between the location of the utterance of the voice signal (1905) and the location of the speaker (1903).

[0376] Since the wearable device (1301) can identify who the speaker of the voice signal (e.g., voice signal (1810), voice signal (1905)) is, it can provide a function according to the voice signal according to each speaker of the voice signal. For example, the function according to the voice signal can include a function of displaying text representing the voice signal (e.g., text (1927), text (1928)) and / or a function of outputting an audio signal representing the voice signal (e.g., audio signal (1930), audio signal (1931)).

[0377] The wearable device (1301) can determine a speaker. For example, the wearable device (1301) can determine a speaker (1802) among speakers by receiving a user input (1922) for a visual object (1921) within a screen (1920). For example, the user input (1922) for the visual object (1921) can change from example (1901) to example (1902). For example, the user input (1922) for the visual object (1921) can include at least one of a gesture input for the visual object (1921), a gaze input for the visual object (1921), or a voice input for the visual object (1921).

[0378] In example (1902), the wearable device (1301) can continue to provide a function according to a voice signal (1810) in response to a user input (1922) for a visual object (1921). The wearable device (1301) can continue to display text (1927) in an area (1925) within the screen (1920). The wearable device (1301) can continue to output an audio signal (1930) through a speaker (1355).

[0379] The wearable device (1301) may stop, skip, refrain from, or bypass providing a function according to a voice signal (1905) based on a user input (1922) to a visual object (1921). For example, the wearable device (1301) may stop, skip, refrain from, or bypass displaying text (1928) based on a user input (1922) to a visual object (1921). For example, the text (1928) may not be displayed on the screen (1920). The wearable device (1301) may stop, skip, refrain from, or bypass outputting an audio signal (1931) through a speaker (1355).

[0380] According to one embodiment, the wearable device (1301) may perform cancellation on a voice signal (1905) acquired through a microphone (1365) according to a user input (1922) for a visual object (1921). For example, the wearable device (1301) may output an audio signal (1930) among the audio signal (1931) and the audio signal (1930) according to a user input (1922) for a visual object (1921).

[0381] As a non-limiting example, the wearable device (1301) may output a voice signal (1905) and an audio signal (1930) acquired through a microphone (1365) through a speaker (1355) in response to a user input (1922) for a visual object (1921).

[0382] In one embodiment, the wearable device (1301) can track the visual object (1921) based on a user input (1922) for the visual object (1921). For example, the wearable device (1301) can identify whether the visual object (1921) is positioned within the screen (1920) based on the user input (1922) for the visual object (1921). For example, the wearable device (1301) can continue to provide the function according to the voice signal (1810) while the visual object (1921) is positioned within the screen (1920). For example, the wearable device (1301) can stop providing the function according to the voice signal (1810) based on identifying that at least a portion of the visual object (1921) is not positioned within the screen (1920). For example, the wearable device (1301) may temporarily disable providing functionality in response to a voice signal (1810).

[0383] For example, the wearable device (1301) may activate a timer based on identifying that at least a portion of a visual object (1921) is not positioned within the screen (1920). For example, the wearable device (1301) may provide a function according to the voice signal (1810) based on identifying that the visual object (1921) is positioned within the screen (1920) before the timer expires. For example, the wearable device (1301) may terminate the function according to the voice signal (1810) in response to the expiration of the timer. As a non-limiting example, the wearable device (1301) may not terminate the function according to the voice signal (1810) based on identifying that the visual object (1921) is included within one or more images acquired via the image sensor (1621).

[0384] Although not illustrated in FIG. 19, according to one embodiment, the wearable device (1301) may display a visual effect for a visual object (1921) through a screen (1920) in response to a user input (1922) for the visual object (1921). For example, the visual effect may be used to distinguish the visual object (1921) that is the target of the user input (1922) from other visual objects.

[0385] FIGS. 20A and 20B illustrate examples of a wearable device that adaptively provides responses to audio signals based on a user's gaze. The operations of the wearable device (1301) illustrated in FIGS. 20A and 20B may be executed, performed, or controlled by a processor (e.g., processor (1610)). Any details that overlap with the descriptions of FIGS. 18 and 19 may be omitted.

[0386] Referring to FIG. 20A, in example (2001), a wearable device (1301) can provide a pass-through function to a user (1801). The wearable device (1301) can obtain an audio signal via a microphone (e.g., microphone (1365)). For example, the audio signal can include a voice signal (1810), a voice signal (2005), and / or noise. The wearable device (1301) can obtain sensing data via at least one sensor (e.g., at least one sensor (1620)). For example, the sensing data can include an image obtained via an image sensor (e.g., image sensor (1621)). For example, the image can be an image of an external environment (or physical environment) captured. For example, the image may include a visual object (2021) corresponding to the speaker (1802) and / or a visual object (2023) corresponding to the speaker (2003). For example, the image may include a portion of a visual object (2021) corresponding to the lips of the speaker (1802).

[0387] The wearable device (1301) can provide audio signals acquired via a microphone (1365) and / or sensing data via at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can identify a voice signal (1810) and / or a voice signal (2005) using the first multimodal model (350) even in an environment where the SNR of the audio signal is relatively low.

[0388] In one embodiment, in example (2001), the user (1801) and the speaker (1802) may be looking at each other. For example, the wearable device (1301) may identify the direction of the gaze of the user (1801). For example, the wearable device (1301) may identify the direction of the gaze of the speaker (1802). For example, the wearable device (1301) may identify the direction of the gaze of a visual object (2021) included in an image acquired through the image sensor (1621). For example, the direction of the gaze of the speaker (1802) may correspond to the direction of the gaze of the visual object (2021). For example, the wearable device (1301) may identify the direction of the gaze of the speaker (1802) using the direction of the gaze of the visual object (2021). The wearable device (1301) can identify that the user (1801) and the speaker (1802) are looking at each other based on the direction of the user's (1801) gaze and the direction of the speaker's (1802) gaze. The wearable device (1301) can identify that the user's (1801) gaze direction corresponds to the speaker's (1802) gaze direction, thereby identifying that the user's (1801) gaze direction corresponds to the speaker's (1802) gaze direction. The wearable device (1301) can provide a function according to the voice signal (1810) of the speaker (1802) based on identifying that the user's (1801) gaze direction corresponds to the speaker's (1802) gaze direction. For example, for the function according to the voice signal (1810), reference may be made to the descriptions of FIG. 19. For example, a function according to a voice signal (1810) may include a function of displaying text (2027) representing the voice signal (1810) through an area (2025) within a screen (2020) and / or a function of outputting an audio signal (2030) representing the voice signal (1810).

[0389] In one embodiment, the user (1801) and the speaker (2003) may not be looking at each other. For example, the wearable device (1301) may identify the direction of the gaze of the user (1801). For example, the wearable device (1301) may identify the direction of the gaze of the speaker (2003). For example, the wearable device (1301) may identify the direction of the gaze of a visual object (2023) included in an image acquired through the image sensor (1621). For example, the direction of the gaze of the speaker (2003) may correspond to the direction of the gaze of the visual object (2023). For example, the wearable device (1301) may identify the direction of the gaze of the speaker (2003) using the direction of the gaze of the visual object (2023). The wearable device (1301) can identify that the user (1801) and the speaker (2003) are not looking at each other based on the direction of the user's (1801) gaze and the direction of the speaker's (2003) gaze. As the wearable device (1301) identifies that the direction of the user's (1801) gaze does not correspond to the direction of the speaker's (2003) gaze, the wearable device (1301) can identify that the user (1801) and the speaker (2003) are not looking at each other. Based on identifying that the user (1801) and the speaker (2003) are not looking at each other, the wearable device (1301) can refrain from, stop, skip, or bypass providing a function according to the speaker's (2003) voice signal (1810).

[0390] In one embodiment, depending on a change in the posture of the speaker (2003), the example (2001) may change to the example (2002). In the example (2002), the user (1801) and the speaker (2003) may be looking at each other. For example, the wearable device (1301) may identify the direction of the gaze of the user (1801). For example, the wearable device (1301) may identify the direction of the gaze of the speaker (2003). For example, the wearable device (1301) may identify the direction of the gaze of a visual object (2023) included in an image acquired through the image sensor (1621). For example, the direction of the gaze of the speaker (2003) may correspond to the direction of the gaze of the visual object (2023). For example, the wearable device (1301) can identify the direction of the gaze of the speaker (2003) using the direction of the gaze of the visual object (2023). Based on the direction of the gaze of the user (1801) and the direction of the gaze of the speaker (2003), the wearable device (1301) can identify that the user (1801) and the speaker (2003) are looking at each other. As the wearable device (1301) identifies that the direction of the gaze of the user (1801) corresponds to the direction of the gaze of the speaker (2003), the wearable device can identify that the user (1801) and the speaker (2003) are looking at each other. The wearable device (1301) can provide a function based on the voice signal (1810) of the speaker (2003) based on identifying that the user (1801) and the speaker (2003) are looking at each other. For example, the wearable device (1301) can further display text (2028) in an area (2026) within the screen (2020). For example, the text (2028) can represent the voice signal (2005). For example, the wearable device (1301) can output an audio signal (2040) through the speaker (1355). For example, the audio signal (2040) can represent the voice signal (1810) and / or the voice signal (2005).For example, the audio signal (2040) may correspond to a signal that is a combination of the audio signal (2030) and the voice signal (2005).

[0391] Referring to Fig. 20b, the operations illustrated in Fig. 20b may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0392] In operation 2051, the wearable device (1301) may display a screen representing at least a portion of a three-dimensional space through a display (e.g., display (1350)). For example, the three-dimensional space may be provided according to a pass-through function of the wearable device (1301). For example, the three-dimensional space may correspond to a physical environment. For example, the wearable device (1301) may display a screen representing at least a portion of the three-dimensional space using an image acquired through an image sensor (e.g., image sensor (1361)). For example, the screen representing at least a portion of the three-dimensional space may include one or more visual objects. For example, each visual object within the one or more visual objects may correspond to a speaker.

[0393] In operation 2053, the wearable device (1301) may identify a visual object corresponding to a voice signal included in the audio signal among one or more visual objects included in a screen based on an audio signal acquired through a microphone (e.g., microphone (1365)). For example, the audio signal may be acquired while the screen is displayed through the display (1350). For example, the wearable device (1301) may identify a speaker for the voice signal included in the audio signal. For a method of identifying a speaker for the voice signal included in the audio signal, reference may be made to the descriptions of FIG. 19. For example, the wearable device (1301) may identify a visual object corresponding to a speaker for the voice signal. For example, the wearable device (1301) may identify a speaker corresponding to a visual object among one or more speakers using the audio signal.

[0394] In operation 2055, the wearable device (1301) can identify a first direction of a gaze of a user wearing the wearable device (1301) and a second direction of a gaze of a visual object corresponding to a voice signal, through at least one sensor (e.g., at least one sensor (1620)). For example, the wearable device (1301) can identify the gaze of the user through an image sensor (1621) (e.g., a gaze tracking camera (1360-1)). For example, the gaze of a visual object corresponding to a voice signal can be related to the gaze of a speaker corresponding to the visual object. For example, the gaze of the visual object can correspond to the gaze of the speaker.

[0395] In operation 2057, the wearable device (1301) can identify whether the user wearing the wearable device (1301) and the speaker corresponding to the visual object are looking at each other based on a first direction of the gaze of the user wearing the wearable device (1301) and a second direction of the gaze of the visual object. For example, the wearable device (1301) can identify the location of the user wearing the wearable device (1301) using at least one sensor (1620). For example, the wearable device (1301) can identify the location of the speaker corresponding to the visual object using at least one sensor (1620). For example, the wearable device (1301) can identify whether a first condition that the speaker is located in the first direction from the location of the user and a second condition that the user is located in the second direction from the location of the speaker are each satisfied. For example, the first condition may be referenced as the user's gaze being directed toward the speaker. For example, the second condition may be referenced as the speaker's gaze being directed toward the user.

[0396] The wearable device (1301) can identify that the user and the speaker are looking at each other based on the fulfillment of the first condition and the fulfillment of the second condition. The wearable device (1301) can identify that the user and the speaker are not looking at each other based on the failure (or non-fulfillment) of the first condition and / or the failure (or non-fulfillment) of the second condition. For example, the wearable device (1301) can execute operation 2059 based on identifying that the user and the speaker corresponding to the visual object are looking at each other. For example, the wearable device (1301) can execute operation 2061 based on identifying that the user and the speaker corresponding to the visual object are not looking at each other.

[0397] In operation 2059, the wearable device (1301) may display text generated from the speaker's voice signal, superimposed on the screen through the display (1350), based on identifying that the user and the speaker corresponding to the visual object are looking at each other. For example, the text may represent the voice signal. For example, the text generated from the voice signal may include text converted from the voice signal. For example, the wearable device (1301) may generate the text by applying STT (speech to text) to the voice signal. For example, the text may be displayed in conjunction with a visual object. For example, the text may be displayed next to the visual object.

[0398] As a non-limiting example, the wearable device (1301) may output an audio signal representing a voice signal through a speaker (e.g., speaker (1355)) based on identifying that a user and a speaker corresponding to a visual object are looking at each other.

[0399] In operation 2061, the wearable device (1301) may refrain from, stop, skip, or bypass displaying text generated from the speaker's voice signal through the display (1350) based on the wearable device (1301) identifying that the user and the speaker corresponding to the visual object are not looking at each other. For example, the wearable device (1301) may continue to display the screen of operation 2051 through the display (1350).

[0400] As a non-limiting example, the wearable device (1301) may refrain from, stop, skip, or bypass outputting an audio signal representing a voice signal through the speaker (1355) based on identifying that the user and the speaker corresponding to the visual object are not looking at each other.

[0401] FIG. 21 illustrates an example of a wearable device that provides a response in a second language to an audio signal spoken in a first language. The operations of the wearable device (1301) illustrated in FIG. 21 may be executed, performed, or controlled by a processor (e.g., processor (1610)).

[0402] Referring to FIG. 21, the wearable device (1301) can acquire an audio signal via a microphone (e.g., microphone (1365)). For example, the audio signal can include a voice signal (2110) and / or noise. For example, the voice signal (2110) can be represented by a first language. For example, the voice signal (2110) can include an utterance in the first language. The wearable device (1301) can acquire sensing data via at least one sensor (e.g., at least one sensor (1620)). For example, the sensing data can include an image acquired via an image sensor (e.g., image sensor (1621)). For example, the image can be an image of an external environment (or physical environment) captured. For example, the image can include a visual object (2121) corresponding to a speaker (1802). For example, the image may include a portion of a visual object (2121) corresponding to the lips of the speaker (1802).

[0403] The wearable device (1301) can provide audio signals acquired through a microphone (1365) and / or sensing data through at least one sensor (1620) to a first multimodal model (e.g., the first multimodal model (350)). For example, the wearable device (1301) can identify a voice signal (2110) even in an environment where the SNR of the audio signal is relatively low by using the first multimodal model (350).

[0404] In one embodiment, the first multimodal model (350) may include one or more multimodal models. For example, each of the one or more multimodal models may be trained using a different language. For example, each of the one or more multimodal models may be trained using images representing lips speaking a different language.

[0405] In one embodiment, the wearable device (1301) can identify that the language of the voice signal (2110) is a first language. For example, the wearable device (1301) can determine a multimodal model corresponding to the first language from among one or more multimodal models. For example, if the first language is English, the wearable device (1301) can determine a multimodal model trained using English. For example, if the first language is Arabic, the wearable device (1301) can determine a multimodal model trained using Arabic.

[0406] In one embodiment, the wearable device (1301) can display the screen (2120) through a display (e.g., display (1350)). For example, the wearable device (1301) can display text (2125) corresponding to the voice signal (2110) through an area (2125) within the screen (2120). For example, the text (2125) can be represented by a second language. For example, the context of the voice signal (2110) and the context of the text (2125) can be substantially the same. For example, the text (2125) can be a text representing the voice signal (2110) translated from a first language to a second language. In one embodiment, the wearable device (1301) can output an audio signal (2130) corresponding to the voice signal (2110) through a speaker (e.g., speaker (1355)). For example, the language of the audio signal (2130) may be a second language. For example, the audio signal (2130) may indicate that the voice signal (2110) has been translated from a first language into a second language. For example, the second language may be different from the first language of the voice signal (2110). For example, the second language may be a language set in the wearable device (1301). For example, the second language may be set because the language most frequently spoken by the user (1801) within a given time period is the second language. As a non-limiting example, the second language may be set by the user (1801).

[0407] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains.

[0408] An electronic device as described above may include an image sensor facing a front side of the electronic device. The electronic device may include a microphone. The electronic device may include communication circuitry. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to detect an event for driving the image sensor. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first image through the image sensor driven in response to the event, and to acquire a first audio signal through the microphone, based on the detection. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, a signal to an external electronic device, wherein the signal causes a wake-up of a second multimodal model (e.g., a second multimodal model (360)) within the external electronic device based on identifying, using a first multimodal model (e.g., a first multimodal model (350)) within the electronic device, the first image and the first audio signal including the reference voice command representing the reference gesture of the user. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device, after transmitting the signal, to acquire a second audio signal via the microphone and to acquire a second image via the image sensor.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, to the external electronic device via the communication circuit, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device via the communication circuit, response information regarding the second audio signal, obtained using the second multimodal model while in a wake-up state in response to the signal. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.

[0409] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from the second audio signal among the second image and the second audio signal, a type of environment in which the electronic device is located, using a third model within the electronic device (e.g., the third model (370)). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain data for a third audio signal as the first data through noise canceling on the second audio signal, which is performed by providing the second data, third data for the first type, and fourth data for the second audio signal to a fourth multimodal model (e.g., the fourth multimodal model (380)) based on identifying that the type of the environment is a first type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the fourth data for the second audio signal as the first data based on identifying that the type of the environment is a second type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.

[0410] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information from the external electronic device, the response information representing an emoji graphical object selected by the second multimodal model from among emoji graphical objects available within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.

[0411] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device through the communication circuit, the response information representing a text selected by the second multimodal model from among texts pre-stored in the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.

[0412] According to one embodiment, the electronic device may include a speaker. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output the audio signal through the speaker.

[0413] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, the response information including text generated based on the first data and the second data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text in the response information through the display.

[0414] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second data using the third data.

[0415] A method performed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit as described above may include an operation of detecting an event for driving the image sensor. The method may include an operation of acquiring a first image through the image sensor driven in response to the event based on the detection, and an operation of acquiring a first audio signal through the microphone. The method may include an operation of transmitting, through the communication circuit, a signal to an external electronic device, based on identifying the first image expressing a reference gesture of a user and the first audio signal including a reference voice command, using a first multimodal model within the electronic device (e.g., a first multimodal model (350)), the signal causing a wake-up of a second multimodal model within the external electronic device (e.g., a second multimodal model (360)). After transmitting the signal, the method may include an operation of acquiring a second audio signal through the microphone, and a second image through the image sensor. The method may include an operation of transmitting, to the external electronic device, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips, via the communication circuit. The method may include an operation of receiving, from the external electronic device, response information regarding the second audio signal obtained using the second multimodal model in a wake-up state according to the signal, via the communication circuit. The method may include an operation of performing a function according to the response information.

[0416] According to one embodiment, the method may include an operation of identifying a type of an environment in which the electronic device is located from the second audio signal among the second audio signal and the second image using a third model within the electronic device (e.g., the third model (370)). The method may include an operation of obtaining data for a third audio signal as the first data through noise cancellation of the second audio signal by providing the second data, third data for the first type, and fourth data for the second audio signal to a fourth multimodal model (e.g., the fourth multimodal model (380)) based on identifying that the type of the environment is a first type. The method may include an operation of obtaining the fourth data for the second audio signal as the first data based on identifying that the type of the environment is a second type. The method may further include an operation of transmitting the first data and the second data to the external electronic device through the communication circuit.

[0417] According to one embodiment, the electronic device may include a display. The method may include receiving, from the external electronic device, response information representing an emoji graphical object selected by the second multimodal model from among emoji graphical objects available within the electronic device, via the communication circuit. The method may include performing the function according to the response information by displaying the emoji graphical object through the display.

[0418] According to one embodiment, the electronic device may include a display. The method may include receiving, from the external electronic device, through the communication circuit, response information indicating a text selected by the second multimodal model from among texts pre-stored within the electronic device. The method may include performing the function according to the response information by displaying the text indicated by the response information through the display.

[0419] According to one embodiment, the electronic device may include a speaker. The method may include an operation of obtaining an audio signal by applying text-to-speech (TTS) to the received response information. The method may include an operation of outputting the audio signal through the speaker.

[0420] According to one embodiment, the electronic device may include a display. The method may include an operation of receiving, via the communication circuit, the response information, which includes text generated based on the first data and the second data. The method may include an operation of performing the function according to the response information by displaying the text within the response information through the display.

[0421] In one embodiment, the method may include an operation of obtaining third data for an area including another visual object corresponding to the user's face within the second image using a face detection model within the electronic device. The method may include an operation of obtaining the second data using the third data.

[0422] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and a communication circuit, cause the electronic device to detect an event to drive the image sensor. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to acquire a first image through the image sensor driven according to the event, and to acquire a first audio signal through the microphone, based on the detection. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, through the communication circuit, a signal to the external electronic device, wherein the signal causes a wake-up of a second multimodal model (e.g., a second multimodal model (360)) within the external electronic device based on identifying, using a first multimodal model (e.g., a first multimodal model (350)) within the electronic device, the first image representing the user's reference gesture and the first audio signal including the reference voice command. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to, after transmitting the signal, acquire a second audio signal via the microphone and to acquire a second image via the image sensor. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit, via the communication circuit, first data obtained using the second audio signal and second data regarding a visual object in the second image corresponding to the user's lips, to the external electronic device.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, response information for the second audio signal obtained using the second multimodal model in a wake-up state in response to the signal from the external electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform a function in response to the response information.

[0423] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify a type of an environment in which the electronic device is located from the second audio signal among the second image and the second audio signal using a third model within the electronic device (e.g., the third model (370)). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain data for a third audio signal as the first data through noise canceling on the second audio signal, which is performed by providing the second data, third data for the first type, and fourth data for the second audio signal to a fourth multimodal model (e.g., the fourth multimodal model (380)) based on identifying that the type of the environment is a first type, and obtain the fourth data for the second audio signal as the first data based on identifying that the type of the environment is a second type. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.

[0424] According to one embodiment, the electronic device may include a display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, through the communication circuit, from an external electronic device, response information representing an emoji graphical object selected by the second multimodal model from among emoji graphical objects available within the electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.

[0425] According to one embodiment, the electronic device may include a display. The one or more programs may include instructions that cause the electronic device to receive, through the communication circuit, from the external electronic device, the response information indicating a text selected by the second multimodal model from among texts pre-stored in the electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text indicated by the response information through the display.

[0426] According to one embodiment, the electronic device may include a speaker. The one or more programs may include instructions that cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to output the audio signal through the speaker.

[0427] According to one embodiment, the electronic device may include a display. The one or more programs may include instructions that cause the electronic device to receive, through the communication circuit, the response information including text generated based on the first data and the second data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform the function according to the response information by displaying the text in the response information through the display.

[0428] In one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the second data using the third data.

[0429] As described above, the electronic device may include an image sensor facing a front side of the electronic device. The electronic device may include a microphone. The electronic device may include communication circuitry. The electronic device may include a memory storing instructions and including one or more storage media. The electronic device may include at least one processor including processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first audio signal via the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a type of environment in which the electronic device is located from the first audio signal using a model within the electronic device (e.g., the third model (370)). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, through the communication circuitry, a signal to the external electronic device, wherein the signal causes a wake-up of a first multimodal model (e.g., a second multimodal model (360)) within the external electronic device, in accordance with a reference voice command included in the first audio signal, based on identifying that the type of the environment is a first type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor, based on identifying that the type of the environment is a second type.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a first image via the driven image sensor. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire a second audio signal via the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuit, to the external electronic device, the signal causing a wake-up of the first multimodal model based on identifying, using a second multimodal model (e.g., the first multimodal model (350)) within the electronic device, the first image representing a reference gesture of the user and the second audio signal including the reference voice command.

[0430] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to activate a timer associated with the image sensor based on identifying that the type of the environment is the second type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire the first image via the driven image sensor prior to expiration of the timer based on identifying that the type of the environment is the second type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, the signal to cause a wake-up of the first multimodal model based on identifying the first image representing the reference gesture and the second audio signal including the reference voice command using the second multimodal model, based on identifying the type of the environment to be the second type, and to continue driving the image sensor independently of the expiration of the timer. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to stop driving the image sensor in response to the expiration of the timer, based on identifying the type of the environment to be the second type, and based on identifying the first image not representing the reference gesture and / or the second audio signal not including the reference voice command using the second multimodal model.

[0431] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the signal based on identifying that the type of the environment is the second type, and then acquire a third audio signal via the microphone and acquire a second image via the image sensor. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, to the external electronic device via the communication circuit, first data acquired using the third audio signal and second data regarding a visual object in the second image corresponding to the user's lips. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device via the communication circuit, response information for the third audio signal, the first data acquired using the first multimodal model in a wake-up state in response to the signal. The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.

[0432] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, using the model within the electronic device, the type of the environment in which the electronic device is located from the third audio signal among the third audio signal and the second image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third data for the third audio signal as the first data based on identifying that the type of the environment is the first type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain data for a fourth audio signal as the first data by providing the second data, fourth data for the second type, and the third data to a third multimodal model (e.g., fourth multimodal model (380)) based on identifying that the type of the environment is the second type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the first data and the second data to the external electronic device via the communication circuit.

[0433] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information from the external electronic device, the response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.

[0434] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device through the communication circuit, response information representing text selected by the first multimodal model from among texts pre-stored in the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.

[0435] In one embodiment, the electronic device may include a speaker. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain an audio signal by applying text-to-speech (TTS) to the received response information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output the audio signal through the speaker.

[0436] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, the response information including text generated based on the first data and the second data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text in the response information through the display.

[0437] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third data for an area including another visual object corresponding to the user's face in the second image using a face detection model within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second data using the third data.

[0438] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a third audio signal via the microphone after transmitting the signal based on identifying the type of the environment as the first type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, using the model within the electronic device, from the third audio signal, the type of the environment in which the electronic device is located. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit, via the communication circuitry, data obtained from the third audio signal to the external electronic device based on identifying, from the third audio signal, the type of the environment as the first type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information for the third audio signal, obtained using the first multimodal model while in a wake-up state in response to the signal, from the external electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.

[0439] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to drive the image sensor based on identifying, from the third audio signal, that the type of the environment is the second type, acquire a second image through the driven image sensor, acquire a fourth audio signal through the microphone, and obtain fourth data for a fifth audio signal through noise canceling on the fourth audio signal by providing first data for a visual object in the second image corresponding to the user's lips, second data for the second type, and third data for the fourth audio signal to a third multimodal model (e.g., a fourth multimodal model (380)), and transmit the first data and the fourth data to the external electronic device through the communication circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information for the fourth audio signal, obtained using the first multimodal model while in a wake-up state in response to the signal, from the external electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform a function according to the response information.

[0440] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, response information from the external electronic device, the response information representing an emoji graphical object selected by the first multimodal model from among emoji graphical objects available within the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the emoji graphical object through the display.

[0441] According to one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, from the external electronic device through the communication circuit, response information representing text selected by the first multimodal model from among texts pre-stored in the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform the function according to the response information by displaying the text represented by the response information through the display.

[0442] A method performed by an electronic device having an image sensor facing a front side of the electronic device, a microphone, and communication circuitry as described above may include an operation of acquiring a first audio signal via the microphone. The method may include an operation of identifying a type of an environment in which the electronic device is located from the first audio signal using a model within the electronic device (e.g., a third model (370)). The method may include an operation of transmitting, via the communication circuitry, a signal to an external electronic device, the signal causing a wake-up of a first multimodal model (e.g., a second multimodal model (360)) within the external electronic device, based on identifying that the type of the environment is the first type, according to a reference voice command included in the first audio signal. The method may include an operation of driving the image sensor based on identifying that the type of the environment is the second type. The method may include an operation of acquiring a first image via the driven image sensor. The method may include an action of acquiring a second audio signal through the microphone. The method may include an action of transmitting, to the ext...

Claims

In the electronic device (301), An image sensor (340) facing the front side of the electronic device (301); Microphone (330); Communication circuit (310); A memory (320) including one or more storage media for storing instructions; and At least one processor (300) comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor (300), Detecting an event to drive the image sensor (340), Based on the above detection, a first image is acquired through the image sensor (340) driven according to the event, and a first audio signal is acquired through the microphone (330). Based on identifying the first image expressing the user's reference gesture and the first audio signal including the reference voice command using the first multimodal model (350) in the electronic device (301), a signal causing a wake-up of the second multimodal model (360) in the external electronic device (302) is transmitted to the external electronic device (302) through the communication circuit (310). After transmitting the above signal, a second audio signal is acquired through the microphone (330), and a second image is acquired through the image sensor (340). Transmitting first data obtained using the second audio signal and second data about a visual object in the second image corresponding to the user's lips to the external electronic device (302) through the communication circuit (310), Receiving, from the external electronic device (302), response information for the second audio signal obtained using the second multimodal model (360) in a wake-up state according to the signal, through the communication circuit (310), and To perform functions according to the above response information, causing the above electronic device (301), Electronic device (301). In claim 1, The above instructions, when individually or collectively executed by the at least one processor (300), Using the third model (370) within the electronic device (301), the type of environment in which the electronic device (301) is located is identified from the second audio signal among the second audio signal and the second image, By providing the second data, the third data for the first type, and the fourth data for the second audio signal to a fourth multimodal model (380), based on identifying that the type of the environment is the first type, noise canceling for the second audio signal is performed, thereby obtaining data for the third audio signal as the first data, Based on identifying that the type of the above environment is the second type, acquiring the fourth data for the second audio signal as the first data, and To transmit the first data and the second data to the external electronic device (302) through the communication circuit (310), causing the above electronic device (301), Electronic device (301). In claim 1, Further including a display (311), The above instructions, when individually or collectively executed by the at least one processor (300), Receive, from the external electronic device (302), the response information representing an emoji graphical object (1021) selected by the second multimodal model (360) among the emoji graphical objects available within the electronic device (301), through the communication circuit (310), and To perform the function according to the response information by displaying the emoji graphic object (1021) through the display (311), causing the above electronic device (301), Electronic device (301). In claim 1, Further including a display (311), The above instructions, when individually or collectively executed by the at least one processor (300), Receive, from the external electronic device (302), the response information representing the text (1031) selected by the second multimodal model (360) among the texts pre-stored in the electronic device (301), through the communication circuit (310), and To perform the function according to the response information by displaying the text (1031) indicated by the response information through the display (311), causing the above electronic device (301), Electronic device (301). In claim 1, Including more speakers (312), The above instructions, when individually or collectively executed by the at least one processor (300), Obtaining an audio signal by applying TTS (text to speech) to the above received response information, and To output the audio signal through the speaker (312), causing the above electronic device (301), Electronic device (301). In claim 1, Further including a display (311), The above instructions, when individually or collectively executed by the at least one processor (300), Receive the response information including the text generated based on the first data and the second data through the communication circuit (310), and To perform the function according to the response information by displaying the text in the response information through the display (311), causing the above electronic device (301), Electronic device (301). In claim 1, The above instructions, when individually or collectively executed by the at least one processor (300), Using the face detection model within the electronic device (301), third data is obtained for an area including another visual object corresponding to the user's face within the second image, and To obtain the second data using the third data, causing the above electronic device (301), Electronic device (301). In the electronic device (301), An image sensor (340) facing the front side of the electronic device (301); Microphone (330); Communication circuit (310); A memory (320) storing instructions and including one or more storage media; and At least one processor (300) comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor (300), Through the above microphone (330), a first audio signal is acquired; Using the model (370) within the electronic device (301), the type of environment in which the electronic device (301) is located is identified from the first audio signal; Based on identifying that the type of the above environment is the first type, transmitting a signal to the external electronic device (302) through the communication circuit (310) to cause a wake-up of the first multimodal model (360) within the external electronic device (302) according to a reference voice command included in the first audio signal; and Based on identifying the above type of environment as the second type: Drive the above image sensor (340), Obtaining a first image through the above-mentioned driven image sensor (340), Obtaining a second audio signal through the above microphone (330), and Based on identifying the first image expressing the user's reference gesture and the second audio signal including the reference voice command using the second multimodal model (350) in the electronic device (301), transmit the signal causing wake-up of the first multimodal model to the external electronic device (302) through the communication circuit (310). causing the above electronic device (301), Electronic device (301). In claim 8, The above instructions, when individually or collectively executed by the at least one processor (300), Based on identifying that the above type of the above environment is the second type: Activate the timer associated with the image sensor (340), Before the expiration of the above timer, the first image is acquired through the driven image sensor (340), Based on identifying the first image expressing the reference gesture and the second audio signal including the reference voice command using the second multimodal model (350), transmitting the signal causing wake-up of the first multimodal model to the external electronic device (302) through the communication circuit (310), and maintaining driving the image sensor (340) independently of the expiration of the timer, and Based on identifying the first image that does not express the reference gesture and the second audio signal that does not include the reference voice command using the second multimodal model (350), stop driving the image sensor (340) in response to the expiration of the timer. Causing the above electronic device (301). Electronic device (301). In claim 8, The above instructions, when individually or collectively executed by the at least one processor (300), After transmitting the signal based on identifying that the type of the environment is the second type, a third audio signal is acquired through the microphone (330), and a second image is acquired through the image sensor (340). Transmitting first data obtained using the third audio signal and second data about a visual object in the second image corresponding to the user's lips to the external electronic device (302) through the communication circuit (310), Receiving, from the external electronic device (302), response information for the third audio signal obtained using the first multimodal model (360) in a wake-up state according to the signal, through the communication circuit (310), and To perform functions according to the above response information, causing the above electronic device (301), Electronic device (301). In claim 10, The above instructions, when individually or collectively executed by the at least one processor (300), Using the model (370) within the electronic device (301), the type of the environment in which the electronic device (301) is located is identified from the third audio signal among the third audio signal and the second image, Based on identifying that the type of the above environment is the first type, third data for the third audio signal is acquired as the first data, By providing the second data, the fourth data for the second type, and the third data to a third multimodal model (380), based on identifying that the type of the environment is the second type, noise canceling for the third audio signal is performed, thereby obtaining data for the fourth audio signal as the first data, and To transmit the first data and the second data to the external electronic device (302) through the communication circuit (310), causing the above electronic device (301), Electronic device (301). In claim 10, Further including a display (311), The above instructions, when individually or collectively executed by the at least one processor (300), Receive, from the external electronic device (302), the response information representing an emoji graphical object (1021) selected by the first multimodal model (360) among the emoji graphical objects available within the electronic device (301), through the communication circuit (310), and To perform the function according to the response information by displaying the emoji graphic object (1021) through the display (311), causing the above electronic device (301), Electronic device (301). In claim 10, Further including a display (311), The above instructions, when individually or collectively executed by the at least one processor (300), Receive, from the external electronic device (302), the response information representing the text (1031) selected by the first multimodal model (360) among the texts pre-stored in the electronic device (301), through the communication circuit (310), and To perform the function according to the response information by displaying the text (1031) indicated by the response information through the display (311), causing the above electronic device (301), Electronic device (301). In claim 8, The above instructions, when individually or collectively executed by the at least one processor (300), After transmitting the signal based on identifying that the type of the above environment is the first type, a third audio signal is acquired through the microphone (330), Using the model within the electronic device (301), identifying the type of the environment in which the electronic device (301) is located from the third audio signal, Based on identifying that the type of the environment is the first type from the third audio signal, data obtained from the third audio signal is transmitted to the external electronic device (302) through the communication circuit (310). Receiving, from the external electronic device (302), response information for the third audio signal obtained using the first multimodal model (360) in a wake-up state according to the signal, through the communication circuit (310), and To perform functions according to the above response information, causing the above electronic device (301), Electronic device (301). In claim 14, The above instructions, when individually or collectively executed by the at least one processor (300), Based on identifying from the third audio signal that the type of the environment is the second type: Drive the above image sensor (340), Obtaining a second image through the above-mentioned driven image sensor (340), Obtaining a fourth audio signal through the above microphone (330), By providing first data for a visual object in the second image corresponding to the user's lips, second data for the second type, and third data for the fourth audio signal to a third multimodal model (380), noise canceling for the fourth audio signal is performed, thereby obtaining fourth data for the fifth audio signal, and Transmitting the first data and the fourth data to the external electronic device (302) through the communication circuit (310), Receiving, from the external electronic device (302), response information for the fourth audio signal obtained using the first multimodal model (360) in a wake-up state according to the signal, through the communication circuit (310), and To perform functions according to the above response information, causing the above electronic device (301), Electronic device (301).

Citation Information

Patent Citations

  • A blender

    KR1020210019791A

  • Electromagnetic stirring casting system for semi-solid die casting

    KR1020220136639A

  • Method of applying a solvent-borne coating composition to a substrate utilizing a high transfer efficiency applicator to form a coating layer thereon

    KR1020220141258A

  • Water purifer for creating water of humidifier

    KR102261373B1

  • Speech Recognition With Selective Use Of Dynamic Language Models

    US20170186432A1