Electronic device, method, and non-transitory computer-readable storage medium for identifying audio signal using image representing real environment
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-08-13
Smart Images

Figure KR2026095004_13082026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for identifying audio signals using an image representing a real environment
[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for identifying an audio signal using an image representing an actual environment.
[0002] The electronic device may be worn on the head of the user of the electronic device. The electronic device may include a camera and a display. The camera included in the electronic device may be used to acquire an image of the actual environment containing the electronic device. The electronic device may display the image of the actual environment acquired through the camera through the display.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure.
[0004] No claim or determination is made as to whether any of the foregoing can be applied as prior art related to the present disclosure.
[0005] A head-wearable electronic device is described. The head-wearable electronic device may include a memory comprising one or more storage media for storing instructions. The head-wearable electronic device may include one or more cameras available for acquiring images of the actual environment in front of the head-wearable electronic device. The head-wearable electronic device may include a plurality of microphones. The head-wearable electronic device may include at least one processor comprising a processing circuit. The instructions may cause the head-wearable electronic device to acquire location information of a source object included in the actual environment by using the images acquired through the one or more cameras when executed individually or collectively by the at least one processor. The instructions may cause the head-wearable electronic device to acquire audio signals generated within the actual environment and characteristic information related to the audio signals through the plurality of microphones when acquiring the location information when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, the head-wearable electronic device may be enabled to identify a target audio signal corresponding to the source object among the audio signals using the location information and the characteristic information.
[0006] A method is provided. The method may be executed within a head-wearable electronic device having one or more cameras and a plurality of microphones available for acquiring images of a real environment in front of the head-wearable electronic device. The method may include an operation of acquiring location information for a source object included in the real environment using the images acquired through the one or more cameras. When acquiring the location information, the method may include an operation of acquiring audio signals generated in the real environment and characteristic information related to the audio signals through the plurality of microphones. The method may include an operation of identifying a target audio signal corresponding to the source object among the audio signals using the location information and the characteristic information.
[0007] A non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the head-wearable electronic device to acquire location information for a source object included in the actual environment by using the images acquired through the one or more cameras when executed by the head-wearable electronic device having one or more cameras and a plurality of microphones available to acquire images of the actual environment in front of the head-wearable electronic device. The one or more programs may include instructions that cause the head-wearable electronic device to acquire audio signals generated within the actual environment and characteristic information related to the audio signals through the plurality of microphones when executed by the head-wearable electronic device to acquire the location information. The above one or more programs may include instructions that cause the head-wearable electronic device to identify, among the audio signals, a target audio signal corresponding to the source object using the location information and the characteristic information when executed by the head-wearable electronic device.
[0008] Figure 1 illustrates an example of a head-wearable electronic device that displays an image acquired through a camera.
[0009] Figure 2 is a simplified block diagram of an exemplary head-wearable electronic device.
[0010] FIG. 3 is a flowchart illustrating the operation of a head-wearable electronic device that identifies a target audio signal using position information and characteristic information.
[0011] FIG. 4 illustrates an exemplary operation of a head-wearable electronic device for determining a candidate source object.
[0012] FIG. 5 illustrates an exemplary operation of a head-wearable electronic device that predicts characteristic information based on position information.
[0013] FIG. 6 illustrates an exemplary operation of a head-wearable electronic device for acquiring characteristic information of audio signals.
[0014] FIG. 7 illustrates an exemplary operation of a head-wearable electronic device displaying a visual representation to distinguish speakers.
[0015] FIG. 8 illustrates an exemplary operation of a head-wearable electronic device that tracks a source object.
[0016] FIG. 9 is a block diagram of an electronic device in a network environment according to various embodiments.
[0017] FIG. 10a shows an example of a perspective view of a wearable device.
[0018] FIG. 10b shows an example of one or more hardware components placed within a wearable device.
[0019] FIGS. 11a and 11b show examples of the appearance of a wearable device.
[0020] Figure 12 shows an example of a block diagram of a wearable device.
[0021] Figure 13 shows an example of a block diagram of an electronic device for displaying an image in virtual space.
[0022] Figure 1 illustrates an example of a head-wearable electronic device that displays an image acquired through a camera.
[0023] The environment (150) may include a head-wearable electronic device (100), a user (120) of the head-wearable electronic device (100), a participant (130), and a participant (140). The head-wearable electronic device (100) may be used to be worn on the head of the user (120). The head-wearable electronic device (100) may be referred to as an electronic device or a wearable device. For example, the environment (150) may be referred to as a real environment. The participants (130, 140) may be described as people located in the environment (150). For example, each of the participants (130, 140) may be referred to as a third party.
[0024] A user (120) wearing a head-wearable electronic device (100) can converse with participants (130, 140). The head-wearable electronic device (100) can receive audio signals generated by the user (120) and participants (130, 140) through multiple microphones (e.g., multiple microphones (204) of FIG. 2). The head-wearable electronic device (100) can identify, among the received audio signals, an audio signal corresponding to either the user (120) or the participants (130, 140). For example, the head-wearable electronic device (100) may be required to individually identify audio signals generated by the user (120) and the participants (130, 140), respectively, in order to provide a conversation record to the user (120). For example, a head-wearable electronic device (100) can provide a conversation record in the form of text based on natural language through a display assembly (e.g., the display assembly (208) of FIG. 2) by performing speech-to-text (STT) on each of the first audio signals generated by the user (120), the second audio signals generated by the participant (130), and the third audio signals generated by the participant (140). The head-wearable electronic device (100) may be required to identify who among the user (120) and the participants (130, 140) is the source object generating the audio signals.
[0025] According to one embodiment, a head-wearable electronic device (100) can identify whether an audio signal is being generated by identifying the shape of each participant's (130, 140) lips. While the participants (130, 140) are wearing masks, it may be difficult for the head-wearable electronic device (100) to identify the source object generating the audio signal based on identifying the shape of the lips. Even if the head-wearable electronic device (100) cannot identify the shape of each participant's (130, 140) lips, it may be required to identify what the source object generating the audio signal is.
[0026] According to one embodiment, a head-wearable electronic device (100) may perform beamforming through a plurality of microphones (e.g., a plurality of microphones (204) of FIG. 2). For example, as the head-wearable electronic device (100) performs beamforming, it may receive an audio signal generated by one of the participants (130, 140) and refrain from or bypass receiving another audio signal generated by another of the participants (130, 140). For example, the head-wearable electronic device (100) may be required to track the participants (130, 140) to continuously perform beamforming.
[0027] According to one embodiment, a head-wearable electronic device (100) can obtain location information for each of the participants (130, 140) by using image(s) of an environment (150) obtained through one or more first cameras (e.g., one or more first cameras (209) of FIG. 2). For example, the head-wearable electronic device (100) can identify the direction of each of the participants (130, 140) by using the images obtained through the one or more first cameras. For example, the head-wearable electronic device (100) can predict a phase difference or a time difference of arrival between audio signals to be obtained based on the identified direction. For example, after predicting the phase difference, the head-wearable electronic device (100) can receive the audio signals through a plurality of microphones (e.g., a plurality of microphones (204) of FIG. 2). For example, the head-wearable electronic device (100) can determine, among the participants (130, 140), the participant causing the received audio signals based on the phase difference identified using the received audio signals and the predicted phase difference. For example, the head-wearable electronic device (100) can determine a target audio signal among the received audio signals.
[0028] For example, a head-wearable electronic device (100) may include hardware components used to perform or execute the above operations. The hardware components are described and illustrated with reference to FIG. 2.
[0029] Figure 2 is a simplified block diagram of an exemplary head-wearable electronic device.
[0030] Referring to FIG. 2, the electronic device (100) may include one or more first cameras (209), one or more second cameras (210), a display assembly (208), a plurality of microphones (204), at least one processor (207), and a memory (206).
[0031] At least one processor (207) may include a hardware component for processing data using instructions stored in memory (206). The hardware component for processing data may include a CPU (central processing unit) (e.g., including processing circuits). The hardware component for processing data may include a GPU (graphic processing unit) (e.g., including processing circuits). The hardware component for processing data may include a DPU (display processing unit) (e.g., including processing circuits). The hardware component for processing data may include a NPU (neural processing unit) (e.g., including processing circuits).
[0032] At least one processor (207) may include one or more cores. For example, at least one processor (207) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0033] Memory (206) may include a hardware component for storing data and / or instructions that are input to and / or output from at least one processor (207). Memory (206) may include, for example, volatile memory such as RAM (random-access memory) and / or non-volatile memory such as ROM (read-only memory). Volatile memory may include, for example, at least one of DRAM (dynamic RAM), SRAM (static RAM), cache RAM, and PSRAM (pseudo SRAM). Non-volatile memory may include, for example, at least one of PROM (programmable ROM), EPROM (erasable PROM), EEPROM (electrically erasable PROM), flash memory, hard disk, compact disk, and EMMC (embedded multimedia card).
[0034] The display assembly (208) can output visualized information. For example, the display assembly (208) can output visualized information to a user under the control of at least one processor (207). The display assembly (208) may include hardware components of an electronic device (100) used to display a screen. For example, the display assembly (208) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light-emitting diode (OLED) or a micro LED. However, it is not limited thereto. For example, the display assembly (208) may include a liquid crystal display (LCD).
[0035] As a non-limiting example, the display assembly (208) may include a first display positioned in front of the left eye of a user wearing the head-wearable electronic device (100) and a second display positioned in front of the right eye of a user wearing the head-wearable electronic device (100). For example, a first content provided through a screen displayed through the first display may be (substantially) identical to a second content provided through a screen displayed through the second display. Although the first content and the second content are identical to each other, the screen displayed through the second display may have disparity with respect to the screen displayed through the first display. For example, the disparity may cause the display assembly (208) to provide content in three dimensions (e.g., corresponding to the first content and the second content).
[0036] One or more first cameras (209) may include one or more light sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate an electrical signal indicating the color and / or brightness of light. For example, one or more first cameras (209) may be described as one or more image sensors. For example, each of one or more first cameras (209) may be available to acquire an image of the environment surrounding the head-wearable electronic device (100). For example, at least some of the one or more first cameras (209) may have a field of view (FOV) corresponding to the field of view (FOV) of the user's eyes. For example, the FOV of some of the one or more first cameras (209) may be different from the FOV of other parts of the one or more first cameras (209). For example, one or more first cameras (209) may be used to identify input means located around a head-wearable electronic device (100).
[0037] One or more second cameras (210) may include one or more light sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate an electrical signal indicating the color and / or brightness of light. For example, one or more second cameras (210) may be described as one or more image sensors. For example, each of one or more second cameras (210) may be available to acquire an image of the user's eye of the head-wearable electronic device (100). For example, at least some of the one or more second cameras (210) may have a field of view (FOV) corresponding to the field of view (FOV) of the user's eye. For example, the FOV of some of the one or more second cameras (210) may be different from the FOV of other parts of the one or more second cameras (210). For example, one or more second cameras (210) may be used to identify input means located around the head-wearable electronic device (100).
[0038] The sensor (211) may be used to measure the distance (or depth value) between the head-wearable electronic device (100) and an external object. For example, the sensor (211) may be used to measure the distance between the head-wearable electronic device (100) and an external object based on a time-of-flight (ToF) technique. For example, the ToF technique may include an indirect time-of-flight (iToF) technique and a direct time-of-flight (dToF) technique. For example, the iToF technique may be described as a technique for measuring distance by analyzing the phase difference of light reflected from an external object. For example, the dToF technique may be described as a technique for measuring distance by analyzing the difference in arrival time of light reflected from an external object. The sensor (211) may not be an essential hardware component. For example, the sensor (211) may be an optional hardware component.
[0039] A plurality of microphones (204) may include hardware components to support the reception of audio (e.g., the voice of the user (120)) or audio signals. The plurality of microphones (204) may be used to acquire audio data or audio signals by receiving audio, voice, speech, or utterance.
[0040] At least one processor (207) can obtain location information (e.g., location information (510) of FIG. 5) for a source object (e.g., source object (605) of FIG. 6) included in the actual environment (150) by using images of the actual environment (150) obtained through one or more first cameras (209). For example, when the at least one processor (207) obtains the location information, it can obtain audio signals generated within the actual environment and characteristic information related to the audio signals (e.g., second characteristic information (540) of FIG. 5) through a plurality of microphones (204). For example, the at least one processor (207) can identify a target audio signal corresponding to the source object by using the location information and the characteristic information. For example, the at least one processor (207) can obtain text corresponding to the target audio signal by performing STT on the target audio signal. For example, at least one processor (207) can display the acquired text through a display assembly (208). For example, when the at least one processor (207) acquires the target audio signal, it can display a visual representation (e.g., the visual representation (722) of FIG. 7) for representing a source object corresponding to the target audio signal through a display assembly (208).
[0041] FIG. 3 is a flowchart illustrating the operation of a head-wearable electronic device that identifies a target audio signal using location information and characteristic information. This method may be executed by the head-wearable electronic device (100) illustrated in FIG. 2 or by at least one processor (207) of the head-wearable electronic device (100). In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0042] In operation 310, at least one processor (207) can obtain location information (e.g., location information (510) of FIG. 5) for a source object (e.g., source object (605) of FIG. 6) included in the actual environment (150) using images obtained through one or more first cameras (209). For example, one or more first cameras (209) may be used to obtain images of the actual environment (150) in front of the head-wearable electronic device (100). For example, the actual environment (150) may include a source object (e.g., source object (605) of FIG. 6). For example, the source object may be described as an object that causes an audio signal. For example, the source object may include a participant (e.g., participant (130) of FIG. 1).
[0043] According to one embodiment, a head-wearable electronic device (100) can identify the location of each source object (e.g., participants (130, 140)) included in the environment (150) using images of the environment (150) obtained through one or more first cameras (209). For example, the images may include visual objects corresponding to the source objects (e.g., visual object (430) and visual object (440) of FIG. 4). For example, the head-wearable electronic device (100) can determine the orientation of the source objects based on the location (or orientation) of the visual objects identified within the images. For example, the head-wearable electronic device (100) may determine the direction of a source object (e.g., participant (130)) corresponding to a visual object (e.g., visual object (430) of FIG. 4) to be the 11 o'clock direction of the head-wearable electronic device (100) based on the determination that a visual object (e.g., visual object (430) of FIG. 4) included in images acquired through one or more first cameras (209) is located at the 11 o'clock direction. The head-wearable electronic device (100) may acquire location information (e.g., location information (510) of FIG. 5) for a source object (e.g., source object (605) of FIG. 6) included in the environment (150) using images acquired through one or more first cameras (209). For example, the location information may indicate a positional relationship (e.g., direction and / or distance) between the source object and the head-wearable electronic device (100). For example, at least one processor (207) can obtain a distance (or depth value) between the head-wearable electronic device (100) and each of the participants (130, 140) via the sensor (211). For example, at least one processor (207) can obtain location information (e.g., location information (510) of FIG. 5) representing the depth value obtained via the sensor (211). For example, the acquisition of said location information will be described later with reference to FIG. 5.
[0044] In operation 320, when at least one processor (207) acquires location information (e.g., location information (510) of FIG. 5), it may acquire audio signals (e.g., audio signals (530) of FIG. 5) generated within the actual environment (150) and characteristic information related to the audio signals (e.g., second characteristic information (540) of FIG. 5) through a plurality of microphones (204). For example, at least one processor (207) may predict characteristic information between the audio signals to be acquired (e.g., first characteristic information (520) of FIG. 5) based on the location information. For example, the characteristic information may represent at least one of a phase difference, a time of arrival difference, and / or a volume difference between the audio signals acquired or to be acquired through the plurality of microphones (204). For example, the prediction of the characteristic information and the acquisition of the characteristic information will be described later with reference to FIG. 5.
[0045] In operation 330, at least one processor (207) can identify a target audio signal corresponding to a source object (e.g., source object (605) in Fig. 6) by using location information (e.g., location information (510) in Fig. 5) and characteristic information (e.g., second characteristic information (540) in Fig. 5). For example, at least one processor (207) can determine or identify a set of target audio signals (or target audio signals) among sets of audio signals (or audio signals obtained through a plurality of microphones (204)) obtained through a plurality of microphones (204) based on a determination that the difference between the first characteristic information (e.g., first characteristic information (520) in Fig. 5) and the second characteristic information (e.g., second characteristic information (540) in Fig. 5) is within a reference range. For example, since the head-wearable electronic device (100) uses location information (e.g., location information (510) of FIG. 5) based on an image (e.g., image (505) of FIG. 5), it can prevent or reduce the misconception that the source object is located close to the head-wearable electronic device (100) depending on the high sound pressure of the audio generated from the source object. For example, the identification of the target audio signal will be described later with reference to FIG. 5.
[0046] According to one embodiment, at least one processor (207) can map or match each set of audio signals obtained through a plurality of microphones (204) for each of the source objects (e.g., participants (130, 140)) around the head-wearable electronic device (100). For example, at least one processor (207) can determine or identify a set of audio signals corresponding to a participant (130) from among the sets of audio signals obtained through the plurality of microphones (204). For example, at least one processor (207) can identify or determine a target audio signal corresponding to one (a) source object (e.g., source object (605) of FIG. 6) from among the audio signals obtained through the plurality of microphones (204). For example, since at least one processor (207) acquires audio signals through multiple microphones (204), multiple audio signals can be acquired from a single source object (e.g., source object (605) of FIG. 6). For example, at least one processor (207) can acquire or receive a set of audio signals corresponding to voice generated by a participant (130) through multiple microphones (204).
[0047] According to one embodiment, at least one processor (207) can obtain a set of beamforming information corresponding to participants (130, 140) included in an environment (150) by using location information (e.g., location information (510) of FIG. 5). For example, at least one processor (207) can identify a first location of a participant (130) through an image (505) and / or a sensor (211). For example, at least one processor (207) can obtain first beamforming information corresponding to the first location by using the first location. For example, the first beamforming information may be described as information for performing beamforming on a plurality of microphones (204) to reduce or eliminate audio signals caused by a source object different from the participant (130) (e.g., participant (140)). For example, at least one processor (207) can identify a second location of the participant (140) through an image (505) and / or a sensor (211). For example, at least one processor (207) can obtain second beamforming information corresponding to the second location using the first location. For example, the second beamforming information may be described as information for performing beamforming on a plurality of microphones (204) to reduce or eliminate audio signals caused by a source object different from the participant (140) (e.g., participant (130)). For example, at least one processor (207) can perform beamforming using a set of beamforming information for each of the sets of audio signals obtained through the plurality of microphones (204) when the speaker cannot be identified among the source objects included in the environment (150) using the image (505) (or based on a decision that the speaker is unknown). For example, at least one processor (207) can obtain a first result by performing beam forming for each of the sets using the first beam forming information.For example, at least one processor (207) can obtain a second result by performing beamforming for each of the sets using the second beamforming information. For example, among the first result and the second result, at least one processor (207) can determine the result in which the audio signal is changed more as beamforming is performed (or the result in which the effect of beamforming is higher) as the result corresponding to the speaker.
[0048] According to one embodiment, operations 310 to 330 performed through a head-wearable electronic device (100) are described, but the embodiment is not limited thereto. For example, operations 310 to 330 may be performed by an electronic device including a smartphone.
[0049] According to one embodiment, at least one processor (207) can identify one or more candidate source objects using images obtained through one or more first cameras (209). For example, a candidate source object may be described as a candidate for a source object that causes a target audio signal among source objects that cause an audio signal. For example, images obtained through one or more first cameras (209) may include visual objects (430, 440) corresponding to participants (130, 140). For example, at least one processor (207) may determine only some of the participants (130, 140) as candidate source objects. For example, audio signals caused by a source object (e.g., participant (140)) different from a candidate source object (e.g., participant (130)) may not be used to match with the source object (e.g., participant (140)). For example, at least one processor (207) may refrain from, bypass, or block audio signal(s) generated from a source object different from the candidate source object from matching the source object. For example, the determination of the candidate source object is described and illustrated in more detail with reference to FIG. 4.
[0050] FIG. 4 illustrates an exemplary operation of a head-wearable electronic device for determining a candidate source object.
[0051] Referring to FIG. 4, the state (410) can be described as a state of adding (or determining) a participant (130) as a candidate source object. For example, at least one processor (207) can identify a visual object (430) corresponding to the participant (130) and a visual object (440) corresponding to the participant (140) using an image (460) obtained through one or more first cameras (209). For example, each of the visual object (430) and the visual object (440) can represent the participant (130) and the participant (140).
[0052] According to one embodiment, at least one processor (207) can identify a participant (130) facing the head-wearable electronic device (100) using an image (460). For example, the head-wearable electronic device (100) can determine that the participant's (130) head (or gaze) is facing (or facing) the head-wearable electronic device (100) by performing object recognition on the image (460). For example, the head-wearable electronic device (100) can determine the participant (130) as a candidate source object based on the determination. For example, based on the determination, the participant (130) and the user (120) may be included in the same group (413). For example, the group may be used to track candidate source objects or to identify a set of target audio signals among a set of audio signals obtained through a plurality of microphones (204).
[0053] According to one embodiment, at least one processor (207) can add a participant (130) to a group (413) based on the position of the user's (120) gaze. For example, an image (460) may be described as an image displayed through a display assembly (208) of a head-wearable electronic device (100). For example, at least one processor (207) may identify the position of the user's (120) gaze using one or more second cameras (210). For example, one or more second cameras (210) may be arranged in relation to the user's (120) eyes and may be used to track the position of the user's (120) gaze. For example, at least one processor (207) may identify the position of the user's (120) gaze (450) through one or more second cameras (210) when displaying an image (460) representing an environment (150) through the display assembly (208). For example, at least one processor (207) can use an image (460) to identify the location of the user (120)’s gaze (450) when a part of a visual object (430) corresponding to the participant (130) (e.g., a part corresponding to the participant (130)’s head or gaze) is directed toward the head-wearable electronic device (100). For example, at least one processor (207) can use an image (460) to identify the participant (130) in a group (413) based on identifying that the participant (130)’s head and / or eyes are directed toward the head-wearable electronic device (100) and that the user (120)’s gaze (450) is located on the visual object (430) corresponding to the participant (130). For example, at least one processor (207) can identify the location of the participant (130) (continuously) or identify a set of target audio signals (or target audio signals) originating from the participant (130) based on a determination that the participant (130) is included in the group (413).
[0054] According to one embodiment, when only the participant (130) among the participant (130) and the participant (140) is included in the group (413), at least one processor (207) may refrain from, bypass, or block tracking the location of the participant (140). Based on the determination that the participant (140) is not included in the group (413), at least one processor (207) may refrain from, bypass, or block identifying the set of audio signals corresponding to the participant (140) among the sets of audio signals obtained through the plurality of microphones (204). At least one processor (207) may track the location of the source object included in the group (413). At least one processor (207) may identify the set of audio signals corresponding to the visual object included in the group (413) among the sets of audio signals obtained through the plurality of microphones (204).
[0055] According to one embodiment, at least one processor (207) may determine whether to include the participant (130) in a group (413) based on identifying the shape (or form) of a part (415) (e.g., lips) of the participant (130). For example, at least one processor (207) may identify the form of a part (415) of the participant (130) based on performing object recognition and / or object classification on an image (460). For example, at least one processor (207) may determine whether to include the participant (130) in a group (413) based on the difference between a first form of the part (415) of the participant (130) identified using a first image and a second form of the part of the participant (130) identified using a second image. For example, the second image may be described as an image acquired after the first image among images acquired through one or more first cameras (209). Although an operation to determine whether to include a participant (130) in a group (413) based on identifying the shape of a part (415) of a participant (130) using images has been described, embodiments are not limited thereto. For example, at least one processor (207) may determine whether to include a participant (130) in a group (413) based on at least one of the direction in which a part of the participant (130) (e.g., head or gaze) is facing, identification of the shape of another part of the participant (130) (e.g., lips), and / or the location of the user's (120) gaze (450).
[0056] For example, at least one processor (207) can identify a part of a body (e.g., a head or face) facing a head-wearable electronic device (100) by performing object recognition on an image (505) (or images), and can identify the gaze of a user (120) through one or more second cameras (210). For example, at least one processor (207) can add the body to a group of one or more candidate source objects based on acquiring audio signals (530) when it identifies that the gaze of the user (120) is located on a visual object (e.g., a visual object (430)) corresponding to the part of the body. For example, at least one processor (207) can identify a target audio signal corresponding to a source object (605) among the one or more candidate source objects by using at least one of location information (510), first characteristic information (520), and second characteristic information (540).
[0057] According to one embodiment, at least one processor (207) can identify the state of a source object (e.g., source object (605) in FIG. 6) using an image (e.g., image (505) in FIG. 5) (or images). For example, based on the determination that the state is a target state for outputting audio, at least one processor (207) can use location information (e.g., location information (510) in FIG. 5) and second characteristic information (e.g., second characteristic information (540) in FIG. 5) to determine the audio signal (or set of audio signals) obtained when identifying the target state as a target audio signal (or set of target audio signals) corresponding to the source object. For example, the target state for outputting audio may include a state in which a part (415) of the participant (130) is moving.
[0058] The state (420) can be described as a state in which the participant (140) is included in a group (423) based on the gaze of the participant (130). For example, the group (423) can be described as an example of the group (413). At least one processor (207) can identify a situation in which the participant (130) and the participant (140) are facing each other using images obtained through one or more first cameras (209). For example, at least one processor (207) can identify the direction in which a part of the participant (130) (e.g., head or gaze) is facing and the direction in which another part of the participant (140) (e.g., head or gaze) is facing using images obtained through one or more first cameras (209). For example, at least one processor (207) may include the participant (140) in a group (423) based on a determination that part of the participant (130) (e.g., head or gaze) is facing the participant (140) and another part of the participant (140) (e.g., head or gaze) is facing the participant (130). For example, at least one processor (207) may determine that the participant (140) is included in the group (423) based on the above determination. For example, the group (423) may be included in the group (413). For example, at least one processor (207) may track a source object (e.g., participant (140)) included in the group (423) or match a set of target audio signals.
[0059] According to one embodiment, at least one processor (207) may exclude a source object from a group (413) based on the satisfaction of a specified condition. For example, at least one processor (207) may exclude a participant (140) from a group (413) based on a determination that, in a group containing participants (130, 140), for a specified time (e.g., 10 minutes), a part of the participant (140) (e.g., head and / or root of the body) does not face the participant (130) or the head-wearable electronic device (100). For example, at least one processor (207) may determine that the participant (140) is not included in the group (413) based on a determination that, for the specified time, the part of the participant (140) faces a direction different from the direction facing the participant (130) or the head-wearable electronic device (100).
[0060] According to one embodiment, at least one processor (207) may identify the relationship or context between words identified through audio signals to determine whether to include a participant in a group. For example, at least one processor (207) may obtain or identify multiple words by performing STT on sets of audio signals obtained through multiple microphones (204). For example, at least one processor (207) may determine whether to include a source object that causes some of the multiple words in a group by identifying the relationship between the multiple words. For example, at least one processor (207) may determine whether to include a source object that causes some of the multiple words in a group based on the context between the multiple words. For example, at least one processor (207) may identify or calculate the similarity between a first embedding vector corresponding to a first word included in the multiple words and a second embedding vector corresponding to a second word included in the multiple words. For example, at least one processor (207) may determine that the first source object from which the first word is output and the second source object from which the second word is output are included in the same group based on a determination that the similarity exceeds a threshold similarity. At least one processor (207) may determine that the first source object and the second source object are not in the same group based on a determination that the similarity is less than a threshold similarity.
[0061] At least one processor (207) can obtain location information (e.g., location information (510) of FIG. 5) using an image (e.g., image (505) of FIG. 5). For example, at least one processor (207) can determine or identify the location of a candidate source object included in the environment (150) using the location information. For example, at least one processor (207) can predict first characteristic information (e.g., first characteristic information (520) of FIG. 5) to be obtained through a plurality of microphones (204) based on the location information. For example, after obtaining the first characteristic information, at least one processor (207) can receive or obtain audio signals (e.g., audio signals (530) of FIG. 5) through a plurality of microphones (204). For example, at least one processor (207) can obtain second characteristic information (e.g., second characteristic information (540) of FIG. 5) using the audio signals. For example, at least one processor (207) can match the audio signals with one of the source objects included in the environment (150) based on the first characteristic information and the second characteristic information. For example, the acquisition of the first characteristic information and the second characteristic information is described and illustrated in more detail with reference to FIG. 5.
[0062] FIG. 5 illustrates an exemplary operation of a head-wearable electronic device that predicts characteristic information based on position information.
[0063] Referring to FIG. 5, at least one processor (207) can acquire an image (505) representing an environment (150) through one or more first cameras (209). At least one processor (207) can use the image (505) to acquire location information (510) representing the location of source objects (e.g., participants (130, 140) of FIG. 1) included in the environment (150). For example, the location information (510) may represent the location (e.g., direction and / or distance from a head-wearable electronic device (100)) of a source object (or candidate source object) included in the environment (150). For example, at least one processor (207) can acquire location information (510) representing the location of a source object by performing at least one of object recognition, object classification, and / or image segmentation on the image (505). However, embodiments of the present disclosure are not limited thereto. For example, at least one processor (207) may obtain location information (510) further including the distance (or depth value) between a source object and a head-wearable electronic device (100) through a sensor (211). For example, the location information (510) may indicate at least one of the orientation of a source object (or candidate source object) included in the environment (150) and the depth value between the source object and the head-wearable electronic device (100).
[0064] According to one embodiment, at least one processor (207) can obtain or predict first characteristic information (520) using location information (510). For example, at least one processor (207) can predict or obtain first characteristic information (520) related to at least a portion of audio signals (530) to be obtained through a plurality of microphones (204) based on location information (510). For example, the first characteristic information (520) may represent at least one of a phase difference between audio signals to be received through a plurality of microphones (204), a difference in arrival time between the audio signals, and a difference in volume between the audio signals. For example, the first characteristic information (520) can be described as information in which at least one processor (207) predicts at least one of the phase difference between the audio signals, the arrival time difference between the audio signals, and the volume difference between the audio signals using position information (510) before acquiring audio signals through a plurality of microphones (204).
[0065] For example, a plurality of microphones (204) may include a first microphone and a second microphone. For example, the first microphone may receive a first audio signal generated from a single source object (e.g., participant (130) of FIG. 1). For example, the second microphone may receive a second audio signal generated from the source object. For example, because the first microphone is spaced apart from the second microphone, the first phase of the first audio signal obtained through the first microphone may differ from the second phase of the second audio signal obtained through the second microphone. For example, at least one processor (207) may obtain or predict first characteristic information (520) indicating the difference between the first phase and the second phase using location information (510). However, embodiments of the present disclosure are not limited thereto. For example, since the first microphone is spaced apart from the second microphone, the first sound pressure resulting from the first audio signal obtained through the first microphone may be different from the second sound pressure resulting from the second audio signal obtained through the second microphone. For example, the difference between the first phase and the second phase will be described later with reference to FIG. 6.
[0066] According to one embodiment, the head-wearable electronic device (100) may further include a sensor (211) (e.g., the depth sensor (1130) of FIG. 11b). For example, the sensor may be used to obtain a depth value between the head-wearable electronic device (100) and a source object. For example, at least one processor (207) may use the sensor (211) to obtain or identify the distance between the source object (e.g., a participant (130)) included in the environment (150) and the head-wearable electronic device (100). For example, location information (510) may further include the distance (or depth value) between the source object and the head-wearable electronic device (100) obtained through the sensor (211).
[0067] According to one embodiment, a head-wearable electronic device (100) can predict or obtain first characteristic information (520) by using location information (510). For example, at least one processor (207) can predict at least one of a phase difference, a time of arrival difference, and / or a volume difference between audio signals to be generated from said source object by using the location of a source object identified through the location information (510).
[0068] According to one embodiment, at least one processor (207) may receive or acquire audio signals (530) through a plurality of microphones (204). For example, the audio signals (530) may be generated from a source object (or candidate source object) included in the environment (150). At least one processor (207) may acquire or identify second characteristic information (540) using the audio signals (530). For example, at least one processor (207) may acquire second characteristic information (540) indicating at least one of a phase difference, a time of arrival difference, and / or a volume difference between audio signals received through the plurality of microphones (204). For example, the second characteristic information (540) may be described as information indicating at least one of a phase difference between audio signals acquired through the plurality of microphones (204), a time of arrival difference between the audio signals, and / or a volume difference between the audio signals.
[0069] According to one embodiment, at least one processor (207) can identify a difference between first characteristic information (520) predicted (or acquired) based on location information (510) and second characteristic information (540) acquired based on audio signals (530). For example, at least one processor (207) can match the audio signals (530) with a candidate source object (or source object) that causes the audio signals (530) based on a determination that the difference is within a reference range. For example, at least one processor (207) can identify or determine a set of target audio signals corresponding to the candidate source object from among sets of audio signals acquired through a plurality of microphones (204) based on a determination that the difference is within a reference range. For example, at least one processor (207) may determine that at least a portion of the audio signals (e.g., audio signals (530)) to be acquired through a plurality of microphones (204) corresponds to a source object (e.g., source object (605) of FIG. 6) based on the determination that the difference between the first characteristic information (520) and the second characteristic information (540) is within a reference range.
[0070] According to one embodiment, at least one processor (207) can identify one or more candidate source objects using at least one of the depth values obtained through an image (505) and a sensor (211). At least one processor (207) can match or map one of the one or more candidate source objects to one of the sets of audio signals obtained through a plurality of microphones (204). For example, the set of audio signals may be generated from a single source object (or candidate source object). For example, audio generated from a single source object may be converted into a set of audio signals through a plurality of microphones (204). For example, each of the first microphone included in the plurality of microphones (204) and the second microphone included in the plurality of microphones (204) may receive the first audio signal and the second audio signal converted from the audio generated from the single source object. For example, the first phase of the first audio signal and the second phase of the second audio signal may be different from each other. For example, the difference between the first phase and the second phase is explained and illustrated in more detail with reference to FIG. 6.
[0071] FIG. 6 illustrates an exemplary operation of a head-wearable electronic device for acquiring characteristic information of audio signals.
[0072] Referring to FIG. 6, the head-wearable electronic device (100) may include a plurality of microphones (204). For example, the plurality of microphones (204) may include a first microphone (610) and a second microphone (620). The source object (605) may be described as an object that is included in an environment (150) containing the head-wearable electronic device (100) and causes audio (or voice). For example, the source object (605) may include a participant (130) or a participant (140). However, it is not limited thereto. For example, the source object (605) may include an object different from a person (e.g., a smartphone or a television).
[0073] A head-wearable electronic device (100) can acquire or receive audio generated from a source object (605) through a plurality of microphones (204). For example, the head-wearable electronic device (100) can acquire a first audio signal generated from the source object (605) through a first microphone (610). For example, the head-wearable electronic device (100) can acquire a second audio signal generated from the source object (605) through a second microphone (620). Referring to FIG. 6, two microphones included in the plurality of microphones (204) are shown, but the embodiment is not limited thereto.
[0074] A first audio signal received through the first microphone (610) may be different from a second audio signal received through the second microphone (620). For example, the first phase of the first audio signal received through the first microphone (610) may be different from the second phase of the second audio signal. For example, the first phase may be different from the second phase because the distance between the source object (605) and the first microphone (610) is different from the distance between the source object (605) and the second microphone (620). For example, the first phase may be different from the second phase because a head-wearable electronic device (100) or a user (120) is positioned between the first microphone (610) and the second microphone (620).
[0075] The first time required for audio generated from the source object (605) to be received through the first microphone (610) may be different from the second time required for audio generated from the source object (605) to be received through the second microphone (620). For example, the first time may be different from the second time because the distance between the source object (605) and the first microphone (610) is different from the distance between the source object (605) and the second microphone (620). For example, the first time may be different from the second time because a head-wearable electronic device (100) or a user (120) is positioned between the first microphone (610) and the second microphone (620).
[0076] The first volume level of the first audio signal may be different from the second volume level of the second audio signal. For example, the first volume level may be different from the second volume level because the distance between the source object (605) and the first microphone (610) is different from the distance between the source object (605) and the second microphone (620). For example, the first volume level may be different from the second volume level because a head-wearable electronic device (100) or a user (120) is located between the first microphone (610) and the second microphone (620).
[0077] The first sound pressure resulting from the first audio signal may be different from the second sound pressure resulting from the second audio signal. For example, the first sound pressure may be different from the second sound pressure because the distance between the source object (605) and the first microphone (610) is different from the distance between the source object (605) and the second microphone (620). For example, the first sound pressure may be different from the second sound pressure because a head-wearable electronic device (100) or a user (120) is positioned between the first microphone (610) and the second microphone (620).
[0078] The frequency of audio generated from the source object (605) may be changed due to diffraction that occurs as it passes through the head-wearable electronic device (100) and / or the user (120) of the head-wearable electronic device (100). For example, due to said diffraction, the frequency of the first audio signal received through the first microphone (610) may be different from the frequency of the second audio signal received through the second microphone (620).
[0079] At least one processor (207) can obtain or predict first characteristic information (520) by identifying the location of a source object (605) through location information (510). At least one processor (207) can obtain or identify second characteristic information (540) through audio signals (530) generated from the source object (605). For example, the second characteristic information may include at least one of a phase difference, a time of arrival difference, and / or a volume level difference between audio signals converted from audio generated from the source object (605). For example, at least one processor (207) can determine whether a set of audio signals obtained through a plurality of microphones (204) is a set of audio signals generated from one or more candidate source objects based on the difference between the first characteristic information (520) and the second characteristic information (540). For example, at least one processor (207) can match audio signals (530) obtained through a plurality of microphones (204) with one of one or more candidate source objects. For example, at least one processor (207) can determine that the audio signals (530) correspond to said one.
[0080] At least one processor (207) can match or map one source object among one or more candidate source objects to one set of sets of audio signals obtained through a plurality of microphones (204). For example, at least one processor (207) can obtain text corresponding to the matched audio signals by performing STT on the set of matched audio signals. For example, at least one processor (207) can display a UI (User Interface) object representing the obtained text through a display assembly (208). At least one processor (207) can display a visual representation (e.g., visual representation (722) of FIG. 7) linked to a visual object (430) corresponding to the matched audio signals through a display assembly (208). For example, the display of the UI object and the visual object is described and illustrated in more detail with reference to FIG. 7.
[0081] FIG. 7 illustrates an exemplary operation of a head-wearable electronic device displaying a visual representation to distinguish speakers.
[0082] Referring to FIG. 7, the state (710) can be described as a state in which visual representations (722, 724) for distinguishing speakers and a UI object (726) representing text based on speech are displayed. At least one processor (207) can match at least some of the sets of audio signals obtained through a plurality of microphones (204) to each of one or more candidate source objects (individually). For example, at least one processor (207) can determine or identify a set of target audio signals corresponding to a single candidate source object included in one or more candidate source objects. For example, at least one processor (207) can determine a source object (605) corresponding to a target audio signal among candidate source objects included in the environment (150) identified using an image (505) (or images obtained through a plurality of microphones (204)) by using at least one of location information (510), first characteristic information (520), and second characteristic information (540).
[0083] According to one embodiment, at least one processor (207) can obtain text (or a set of texts) corresponding to sets of audio signals by performing STT on sets of audio signals. For example, at least one processor (207) can obtain a first text representing the sentence "Did you see the new Galaxy?" and a second text representing the sentence "Yes, I'm going to buy one." For example, each of the first text and the second text may be uttered by different speakers. For example, at least one processor (207) can perform STT on sets using a trained model. For example, the trained model may include a Large Language Model (LLM). For example, at least one processor (207) can display a UI object (726) for representing at least some of the sets of obtained texts through a display assembly (208). For example, a set of texts contained within a UI object (726) may be displayed with the speaker distinguished. For example, at least one processor (207) may indicate that the first text of the set is induced by a participant (130). For example, at least one processor (207) may indicate that the second text of the set is induced by a participant (140). For example, at least one processor (207) may refer to the participant (130) corresponding to the first text as a first word (e.g., Circle). For example, at least one processor (207) may refer to the participant (140) corresponding to the second text as a second word (e.g., Triangle). For example, the types of words for identifying participants within the UI object (726) may include shapes, colors, numbers, and / or nicknames.
[0084] According to one embodiment, at least one processor (207) can display visual representations (722, 724) associated with a visual object (430) and a visual object (440) corresponding to each of the participant (130) and the participant (140). For example, at least one processor (207) can identify a target audio signal corresponding to a source object (605) among audio signals obtained through a plurality of microphones (204), and then, based on obtaining audio signals caused by the source object (605) through the plurality of microphones (204), display a visual representation (722) to indicate the source object (605) causing the audio signals through a display assembly (208). For example, at least one processor (207) may display a visual representation (722) around a visual object (430) to display a visual object (430) that matches the first word that appears by the UI object (726). For example, at least one processor (207) may display a visual representation (724) around a visual object (440) to display a visual object (440) that matches the second word that appears by the UI object (726). For example, a user (120) may distinguish the speaker based on the visual representations (722, 724). For example, a user (120) may identify who the speaker is that causes the text displayed on the UI object (726) based on the visual representations (722, 724).
[0085] According to one embodiment, when at least one processor (207) performs STT on sets of audio signals obtained through a plurality of microphones (204), it can match sets of audio signals with source objects by using context between identified words. For example, when it is unclear whether a first word obtained by performing STT on the sets is caused by a first source object or a second source object, at least one processor (207) can determine the source object causing the first word by using context. For example, at least one processor (207) can determine whether the first word is caused by the first source object based on identifying the relationship between the words caused by the first source object and the first word. For example, at least one processor (207) can determine whether the first word is caused by the second source object based on identifying the relationship between other words caused by the second source object and the first word.
[0086] According to one embodiment, at least one processor (207) can identify the state of a part of a speaker (e.g., part (415) of FIG. 4) through one or more first cameras (209) in order to match one of the audio signals obtained through a plurality of microphones (204) with a speaker (e.g., participant (130) or participant (140)). For example, at least one processor (207) can identify whether a part (415) of a participant (e.g., participant (130)) is moving by using images obtained through one or more first cameras (209) (e.g., images including image (505)). For example, at least one processor (207) can further determine that the participant is a speaker corresponding to one of the audio signals based on the determination that the part is moving.
[0087] At least one processor (207) can track or identify one or more candidate source objects included in the actual environment (150) based on images obtained through one or more first cameras (209) (e.g., images (815, 825) of FIG. 8). For example, the tracking of the candidate source objects is described and illustrated in more detail with reference to FIG. 8.
[0088] FIG. 8 illustrates an exemplary operation of a head-wearable electronic device that tracks a source object.
[0089] Referring to FIG. 8, the state (810) can be described as a state in which visual objects (430, 440) corresponding to the participants (130, 140) are displayed. For example, at least one processor (207) can acquire images of the environment (150) through one or more first cameras (209). For example, at least one processor (207) can display the images through a display assembly (208).
[0090] At least one processor (207) can track participants (130, 140) using visual objects (430) and visual objects (440) included in the image (815). For example, at least one processor (207) can track a participant determined as a candidate source object using an image (815) obtained through one or more first cameras (209). For example, at least one processor (207) can track a participant determined as a candidate source object by performing object tracking or human tracking on the images (815, 825). For example, at least one processor (207) can track or identify a participant determined as a candidate source object (continuously) by performing at least one of object tracking and / or face recognition on the images (815, 825).
[0091] A state (820) can be described as a state in which an image (825) containing a visual object (440) among visual objects (430, 440) is displayed. For example, as a participant (140) corresponding to a visual object (440) moves, an image (825) containing only the visual object (440) can be displayed through a display assembly (208). At least one processor (207) can (continuously) track a source object (e.g., participant (130) of FIG. 1) corresponding to a visual object (430) using images (815, 825) when the state of an environment (150) including a head-wearable electronic device (100) transitions from state (810) to state (820). For example, at least one processor (207) may retain data (e.g., an identifier based on location or object recognition) for the source object corresponding to the visual object (430), even though the visual object (430) cannot be identified within the image (825). For example, at least one processor (207) may use the data to determine that the visual object (430) added in the image (815) corresponds to the source object when the state transitions from state (820) to state (810). For example, at least one processor (207) may include the source object corresponding to the added visual object in a group (e.g., group (413)).
[0092] According to one embodiment, at least one processor (207) can determine a set of target audio signals corresponding to a target source object among a set of audio signals without controlling the orientation of the one or more first cameras (209) by tracking an object using images (815, 825) obtained through one or more first cameras (209). For example, at least one processor (207) can obtain identification information for each of the visual objects (430, 440) identified within the images (815, 825) by performing object tracking and / or face recognition on the images (815, 825) obtained through one or more first cameras (209). For example, since at least one processor (207) maintains the identification information, it can identify whether a participant located outside the FOV of one or more first cameras (209) (e.g., participant (130) in state (820)) is located within the FOV of one or more first cameras (209). For example, when a participant who was located outside the FOV of one or more first cameras (209) (e.g., participant (130)) is located within the FOV of one or more first cameras (209), at least one processor (207) can use the identification information to identify a set of target audio signals corresponding to the participant among a set of audio signals.
[0093] FIG. 9 is a block diagram of an electronic device in a network environment according to various embodiments.
[0094] FIG. 9 is a block diagram of an electronic device (901) in a network environment (900) according to various embodiments. Referring to FIG. 9, in the network environment (900), the electronic device (901) may communicate with an electronic device (902) through a first network (998) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (904) or a server (908) through a second network (999) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (901) may communicate with the electronic device (904) through a server (908). According to one embodiment, the electronic device (901) may include a processor (920), memory (930), input module (950), sound output module (955), display module (960), audio module (970), sensor module (976), interface (977), connection terminal (978), haptic module (979), camera module (980), power management module (988), battery (989), communication module (990), subscriber identification module (996), or antenna module (997). In some embodiments, at least one of these components (e.g., connection terminal (978)) may be omitted from the electronic device (901), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (976), camera module (980), or antenna module (997)) may be integrated into a single component (e.g., display module (960)).
[0095] The processor (920) can control at least one other component (e.g., a hardware or software component) of the electronic device (901) connected to the processor (920) by executing software (e.g., a program (940)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (920) can store commands or data received from other components (e.g., a sensor module (976) or a communication module (990)) in volatile memory (932), process the commands or data stored in volatile memory (932), and store the resulting data in non-volatile memory (934). According to one embodiment, the processor (920) may include a main processor (921) (e.g., a central processing unit or an application processor) or an auxiliary processor (923) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (901) includes a main processor (921) and an auxiliary processor (923), the auxiliary processor (923) may be configured to use lower power than the main processor (921) or to be specialized for a designated function. The auxiliary processor (923) may be implemented separately from the main processor (921) or as part thereof.
[0096] The auxiliary processor (923) may control at least some of the functions or states associated with at least one component of the electronic device (901) (e.g., display module (960), sensor module (976), or communication module (990)) on behalf of the main processor (921) while the main processor (921) is in an inactive (e.g., sleep) state, or together with the main processor (921) while the main processor (921) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (923) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (980) or communication module (990)). According to one embodiment, the auxiliary processor (923) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (901) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (908)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0097] The memory (930) can store various data used by at least one component of the electronic device (901) (e.g., processor (920) or sensor module (976)). The data may include, for example, software (e.g., program (940)) and input or output data for related commands. The memory (930) may include volatile memory (932) or non-volatile memory (934).
[0098] The program (940) may be stored as software in memory (930) and may include, for example, an operating system (942), middleware (944), or an application (946).
[0099] The input module (950) can receive commands or data to be used for a component of the electronic device (901) (e.g., processor (920)) from outside the electronic device (901) (e.g., user). The input module (950) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0100] The sound output module (955) can output a sound signal to the outside of the electronic device (901). The sound output module (955) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0101] The display module (960) can visually provide information to an external (e.g., user) of the electronic device (901). The display module (960) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (960) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0102] The audio module (970) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (970) can acquire sound through the input module (950) or output sound through the sound output module (955) or an external electronic device (e.g., electronic device (902)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (901).
[0103] The sensor module (976) can detect the operating state of the electronic device (901) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (976) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0104] The interface (977) may support one or more specified protocols that can be used for the electronic device (901) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (902)). According to one embodiment, the interface (977) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0105] The connection terminal (978) may include a connector through which the electronic device (901) can be physically connected to an external electronic device (e.g., electronic device (902)). According to one embodiment, the connection terminal (978) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0106] The haptic module (979) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (979) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0107] The camera module (980) can capture still images and video. According to one embodiment, the camera module (980) may include one or more lenses, image sensors, image signal processors, or flashes.
[0108] The power management module (988) can manage the power supplied to the electronic device (901). According to one embodiment, the power management module (988) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0109] The battery (989) can supply power to at least one component of the electronic device (901). According to one embodiment, the battery (989) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0110] The communication module (990) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (901) and an external electronic device (e.g., electronic device (902), electronic device (904), or server (908)), and the performance of communication through the established communication channel. The communication module (990) may include one or more communication processors that operate independently of the processor (920) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (990) may include a wireless communication module (992) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (994) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (904) through a first network (998) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (999) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (992) can identify or authenticate the electronic device (901) within a communication network such as the first network (998) or the second network (999) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (996).
[0111] The wireless communication module (992) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (992) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (992) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (992) can support various requirements specified in the electronic device (901), external electronic device (e.g., electronic device (904)), or network system (e.g., second network (999)). According to one embodiment, the wireless communication module (992) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0112] An antenna module (997) can transmit a signal or power to an external source (e.g., an external electronic device) or receive it from an external source. According to one embodiment, the antenna module (997) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (997) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (998) or a second network (999), may be selected from the plurality of antennas, for example, by a communication module (990). A signal or power may be transmitted or received between the communication module (990) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (997).
[0113] According to various embodiments, the antenna module (997) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0114] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0115] According to one embodiment, commands or data may be transmitted or received between the electronic device (901) and an external electronic device (904) through a server (908) connected to a second network (999). Each of the external electronic devices (902, or 904) may be the same or a different type of device as the electronic device (901). According to one embodiment, all or part of the operations performed on the electronic device (901) may be performed on one or more of the external electronic devices (902, 904, or 908). For example, if the electronic device (901) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (901) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (901). The electronic device (901) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (901) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (904) may include an Internet of Things (IoT) device. The server (908) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (904) or the server (908) may be included within a second network (999).The electronic device (901) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0116] FIG. 10a illustrates an example of a perspective view of a wearable device. FIG. 10b illustrates an example of one or more hardware components disposed within the wearable device. According to one embodiment, the wearable device (1099) may have the form of glasses that are wearable on a part of a user's body (e.g., head). The wearable device (1099) of FIG. 10a and FIG. 10b may be an example of the electronic device (901) of FIG. 9. The wearable device (1099) of FIG. 10a and FIG. 10b may be an example of the head-wearable electronic device (100) of FIG. 1. The wearable device (1099) may include a head-mounted display (HMD). For example, the housing of the wearable device (1099) may include a flexible material such as rubber and / or silicone having a shape that adheres to a part of the user's head (e.g., a part of the face covering both eyes). For example, the housing of the wearable device (1099) may include one or more straps that can be twined around the user's head and / or one or more temples that can be attached to the ears of the head.
[0117] Referring to FIG. 10a, a wearable device (1099) according to one embodiment may include at least one display (1050) and a frame (1000) supporting at least one display (1050).
[0118] According to one embodiment, a wearable device (1099) may be worn on a part of a user's body. The wearable device (1099) may provide augmented reality (AR), virtual reality (VR), or mixed reality (MR) that combines augmented reality and virtual reality to a user wearing the wearable device (1099). For example, the wearable device (1099) may display a virtual reality image provided by at least one optical device (1082, 1084) of FIG. 10b on at least one display (1050) in response to a specified gesture of the user obtained through the motion recognition camera (1060-2, 1060-3) of FIG. 10b.
[0119] According to one embodiment, at least one display (1050) can provide visual information to a user. For example, at least one display (1050) may include a transparent or translucent lens. At least one display (1050) may include a first display (1050-1) and / or a second display (1050-2) spaced apart from the first display (1050-1). For example, the first display (1050-1) and the second display (1050-2) may be positioned at locations corresponding to the user's left eye and right eye, respectively.
[0120] Referring to FIG. 10b, at least one display (1050) may provide visual information transmitted from external light to a user through a lens included in at least one display (1050) and other visual information distinct from said visual information. The lens may be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens. For example, at least one display (1050) may include a first surface (1031) and a second surface (1032) opposite to the first surface (1031). A display area may be formed on the second surface (1032) of at least one display (1050). When a user wears the wearable device (1099), external light may be transmitted to the user by being incident on the first surface (1031) and transmitted through the second surface (1032). As another example, at least one display (1050) can display an augmented reality image combined with a virtual reality image provided by at least one optical device (1082, 1084) on a real screen transmitted through external light in a display area formed on a second surface (1032).
[0121] In one embodiment, at least one display (1050) may include at least one waveguide (1033, 1034) that diffracts light emitted from at least one optical device (1082, 1084) and transmits it to a user. At least one waveguide (1033, 1034) may be formed based on at least one of glass, plastic, or polymer. A nano pattern may be formed on the exterior or at least a portion of the interior of at least one waveguide (1033, 1034). The nano pattern may be formed based on a polygonal and / or curved grating structure. Light incident on one end of at least one waveguide (1033, 1034) may be propagated to the other end of at least one waveguide (1033, 1034) by the nano pattern. At least one waveguide (1033, 1034) may include at least one diffractive element (e.g., DOE (diffractive optical element), HOE (holographic optical element)) and at least one reflective element (e.g., a reflective mirror). For example, at least one waveguide (1033, 1034) may be placed within a wearable device (1099) to guide a screen displayed by at least one display (1050) to the user's eye. For example, the screen may be transmitted to the user's eye based on total internal reflection (TIR) occurring within at least one waveguide (1033, 1034).
[0122] A wearable device (1099) can analyze an object included in a real-world image collected through a camera (1060-4), combine a virtual object corresponding to an object among the analyzed objects that is the target of augmented reality provision, and display it on at least one display (1050). The virtual object may include at least one of text and an image regarding various information related to the object included in the real-world image. The wearable device (1099) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the wearable device (1099) can perform spatial recognition (e.g., SLAM (simultaneous localization and mapping)) using a multi-camera and / or time-of-flight (ToF). A user wearing the wearable device (1099) can view the image displayed on at least one display (1050).
[0123] According to one embodiment, the frame (1000) may be formed as a physical structure that allows the wearable device (1099) to be worn on the user's body. According to one embodiment, the frame (1000) may be configured so that when the user wears the wearable device (1099), the first display (1050-1) and the second display (1050-2) can be positioned corresponding to the user's left and right eyes. The frame (1000) may support at least one display (1050). For example, the frame (1000) may support the first display (1050-1) and the second display (1050-2) so that they are positioned corresponding to the user's left and right eyes.
[0124] Referring to FIG. 10a, the frame (1000) may include an area (1020) in which at least a portion comes into contact with a part of the user's body when the user wears the wearable device (1099). For example, the area (1020) of the frame (1000) in contact with a part of the user's body may include an area in contact with a part of the user's nose, a part of the user's ear, and a part of the side of the user's face that the wearable device (1099) comes into contact with. According to one embodiment, the frame (1000) may include a nose pad (1010) that comes into contact with a part of the user's body. When the wearable device (1099) is worn by the user, the nose pad (1010) may come into contact with a part of the user's nose. The frame (1000) may include a first temple (1004) and a second temple (1005) that come into contact with another part of the user's body that is distinct from the part of the user's body.
[0125] For example, the frame (1000) may include a first rim (1001) covering at least a portion of a first display (1050-1), a second rim (1002) covering at least a portion of a second display (1050-2), a bridge (1003) positioned between the first rim (1001) and the second rim (1002), a first pad (1011) positioned along a portion of the edge of the first rim (1001) from one end of the bridge (1003), a second pad (1012) positioned along a portion of the edge of the second rim (1002) from the other end of the bridge (1003), a first temple (1004) extending from the first rim (1001) and fixed to a portion of the wearer's ear, and a second temple (1005) extending from the second rim (1002) and fixed to a portion of the ear opposite to the first. The first pad (1011) and the second pad (1012) may come into contact with a part of the user's nose, and the first temple (1004) and the second temple (1005) may come into contact with a part of the user's face and a part of the ear. The temples (1004, 1005) may be rotatably connected to the rim through the hinge units (1006, 1007) of FIG. 10b. The first temple (1004) may be rotatably connected to the first rim (1001) through a first hinge unit (1006) positioned between the first rim (1001) and the first temple (1004). The second temple (1005) may be rotatably connected to the second rim (1002) through a second hinge unit (1007) disposed between the second rim (1002) and the second temple (1005). According to one embodiment, the wearable device (1099) may identify an external object (e.g., a user's fingertip) touching the frame (1000) and / or a gesture performed by said external object by using a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of the surface of the frame (1000).
[0126] According to one embodiment, the wearable device (1099) may include hardware that performs various functions (e.g., hardware to be described later based on the block diagram of FIG. 2). For example, the hardware may include a battery module (1070), an antenna module (1075), at least one optical device (1082, 1084), speakers (e.g., speakers (1055-1, 1055-2)), a microphone (e.g., microphones (1065-1, 1065-2, 1065-3)), a light-emitting module (not shown), and / or a PCB (printed circuit board) (1090) (e.g., a printed circuit board). The various hardware may be placed within a frame (1000).
[0127] According to one embodiment, a microphone (e.g., microphones (1065-1, 1065-2, 1065-3)) of a wearable device (1099) is positioned on at least a portion of a frame (1000) to acquire a sound signal. A first microphone (1065-1) positioned on a bridge (1003), a second microphone (1065-2) positioned on a second rim (1002), and a third microphone (1065-3) positioned on a first rim (1001) are shown in FIG. 10b, but the number and position of the microphones (1065) are not limited to the embodiment of FIG. 10b. If the number of microphones (1065) included in the wearable device (1099) is two or more, the wearable device (1099) can identify the direction of a sound signal by using a plurality of microphones placed on different parts of the frame (1000).
[0128] According to one embodiment, at least one optical device (1082, 1084) may project a virtual object onto at least one display (1050) to provide various image information to a user. For example, at least one optical device (1082, 1084) may be a projector. At least one optical device (1082, 1084) may be disposed adjacent to at least one display (1050) or included within at least one display (1050) as part of at least one display (1050). According to one embodiment, a wearable device (1099) may include a first optical device (1082) corresponding to a first display (1050-1) and a second optical device (1084) corresponding to a second display (1050-2). For example, at least one optical device (1082, 1084) may include a first optical device (1082) positioned at the edge of a first display (1050-1) and a second optical device (1084) positioned at the edge of a second display (1050-2). The first optical device (1082) may transmit light to a first waveguide (1033) positioned on the first display (1050-1), and the second optical device (1084) may transmit light to a second waveguide (1034) positioned on the second display (1050-2).
[0129] In one embodiment, the camera (1060) may include a shooting camera (1060-4), an eye tracking camera (ET CAM) (1060-1), and / or a motion recognition camera (1060-2, 1060-3). The shooting camera (1060-4), the eye tracking camera (1060-1), and the motion recognition camera (1060-2, 1060-3) may be positioned at different locations on the frame (1000) and may perform different functions. The eye tracking camera (1060-1) may output data indicating the position of the eyes or the gaze of a user wearing the wearable device (1099). For example, the wearable device (1099) may detect the gaze from an image containing the user's pupils obtained through the eye tracking camera (1060-1). A wearable device (1099) can identify an object focused by a user (e.g., a real object, and / or a virtual object) by using the user's gaze acquired through an eye-tracking camera (1060-1). Upon identifying the focused object, the wearable device (1099) can perform a function for interaction between the user and the focused object (e.g., gaze interaction). The wearable device (1099) can represent a portion corresponding to the eyes of an avatar representing the user in a virtual space by using the user's gaze acquired through the eye-tracking camera (1060-1). The wearable device (1099) can render an image (or screen) displayed on at least one display (1050) based on the position of the user's eyes. For example, the visual quality of a first region associated with the gaze within the image and the visual quality of a second region distinct from the first region (e.g., resolution, brightness, saturation, grayscale, PPI) may differ from each other.The wearable device (1099) can acquire an image having a visual quality of a first region and a visual quality of a second region that matches the user's gaze by using foveated rendering. For example, if the wearable device (1099) supports an iris recognition function, user authentication can be performed based on iris information acquired using an eye-tracking camera (1060-1). An example in which the eye-tracking camera (1060-1) is positioned toward the user's right eye is illustrated in FIG. 10b, but the embodiment is not limited thereto, and the eye-tracking camera (1060-1) may be positioned alone toward the user's left eye or toward both eyes.
[0130] In one embodiment, the camera (1060-4) can capture a real image or background to be matched with a virtual image in order to implement augmented reality or mixed reality content. The camera (1060-4) can be used to acquire high-resolution images based on HR (high resolution) or PV (photo video). The camera (1060-4) can capture an image of a specific object located at the position viewed by the user and provide the image to at least one display (1050). The at least one display (1050) can display a single image in which information regarding a real image or background including the image of the specific object acquired using the camera (1060-4) and a virtual image provided through at least one optical device (1082, 1084) are superimposed. The wearable device (1099) can compensate for depth information (e.g., distance between the wearable device (1099) and an external object obtained through a depth sensor) using an image obtained through a shooting camera (1060-4). The wearable device (1099) can perform object recognition using an image obtained through a shooting camera (1060-4). The wearable device (1099) can perform a function of focusing on an object (or subject) within an image (e.g., auto focus) and / or an optical image stabilization (OIS) function (e.g., anti-shake function) using a shooting camera (1060-4). The wearable device (1099) can perform a pass-through function to display an image obtained through a shooting camera (1060-4) superimposed on at least a portion of the screen while displaying a screen representing a virtual space on at least one display (1050). In one embodiment, the camera (1060-4) may be placed on a bridge (1003) positioned between the first rim (1001) and the second rim (1002).
[0131] The eye tracking camera (1060-1) can achieve more realistic augmented reality by tracking the gaze of a user wearing a wearable device (1099), thereby matching the user's gaze with visual information provided to at least one display (1050). For example, when the user looks straight ahead, the wearable device (1099) can naturally display environmental information related to the user's front at the location where the user is situated on at least one display (1050). The eye tracking camera (1060-1) may be configured to capture an image of the user's pupil to determine the user's gaze. For example, the eye tracking camera (1060-1) may receive a gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In one embodiment, the eye tracking camera (1060-1) may be positioned at locations corresponding to the user's left and right eyes. For example, the eye-tracking camera (1060-1) may be positioned within the first rim (1001) and / or the second rim (1002) to face the direction in which the user wearing the wearable device (1099) is located.
[0132] The motion recognition camera (1060-2, 1060-3) can provide a specific event to a screen provided on at least one display (1050) by recognizing the movement of the user's entire body or part thereof, such as the user's torso, hands, or face. The motion recognition camera (1060-2, 1060-3) can recognize the user's gesture, acquire a signal corresponding to the gesture, and provide a display corresponding to the signal to at least one display (1050). The processor can identify the signal corresponding to the gesture and, based on the identification, perform a designated function. The motion recognition camera (1060-2, 1060-3) can be used to perform spatial recognition functions using SLAM and / or depth maps for a 6-degrees-of-freedom pose (6 dof pose). The processor can use the motion recognition camera (1060-2, 1060-3) to perform gesture recognition functions and / or object tracking functions. In one embodiment, a motion recognition camera (1060-2, 1060-3) may be placed on the first rim (1001) and / or the second rim (1002).
[0133] The camera (1060) included within the wearable device (1099) is not limited to the eye-tracking camera (1060-1) and motion recognition camera (1060-2, 1060-3) described above. For example, the wearable device (1099) can identify external objects included within the FoV by using a camera positioned toward the user's FoV. The identification of external objects by the wearable device (1099) can be performed based on a sensor for identifying the distance between the wearable device (1099) and the external object, such as a depth sensor and / or a time of flight (ToF) sensor. The camera (1060) positioned toward the FoV may support an autofocus function and / or an optical image stabilization (OIS) function. For example, the wearable device (1099) may include a camera (1060) (e.g., a face tracking camera) positioned toward the face to acquire an image including the face of a user wearing the wearable device (1099).
[0134] Although not illustrated, according to one embodiment, the wearable device (1099) may further include a light source (e.g., LED) that emits light toward a subject (e.g., user's eye, face, and / or an object outside the FoV) being photographed using a camera (1060). The light source may include an LED of infrared wavelength. The light source may be placed in at least one of the frame (1000) and hinge units (1006, 1007).
[0135] According to one embodiment, the battery module (1070) can supply power to the electronic components of the wearable device (1099). In one embodiment, the battery module (1070) may be placed within the first temple (1004) and / or the second temple (1005). For example, the battery module (1070) may be a plurality of battery modules (1070). The plurality of battery modules (1070) may each be placed in the first temple (1004) and the second temple (1005), respectively. In one embodiment, the battery module (1070) may be placed at the end of the first temple (1004) and / or the second temple (1005).
[0136] The antenna module (1075) can transmit a signal or power to the outside of the wearable device (1099) or receive a signal or power from the outside. In one embodiment, the antenna module (1075) may be placed within the first temple (1004) and / or the second temple (1005). For example, the antenna module (1075) may be placed close to one side of the first temple (1004) and / or the second temple (1005).
[0137] The speaker (1055) can output an acoustic signal to the outside of the wearable device (1099). The acoustic output module may be referred to as the speaker. In one embodiment, the speaker (1055) may be placed within a first temple (1004) and / or a second temple (1005) to be placed adjacent to the ear of a user wearing the wearable device (1099). For example, the speaker (1055) may include a second speaker (1055-2) placed adjacent to the user's left ear by being placed within the first temple (1004), and a first speaker (1055-1) placed adjacent to the user's right ear by being placed within the second temple (1005).
[0138] A light-emitting module (not shown) may include at least one light-emitting element. The light-emitting module may emit light of a color corresponding to a specific state or emit light with an action corresponding to a specific state in order to visually provide information regarding a specific state of the wearable device (1099) to the user. For example, if the wearable device (1099) requires charging, it may emit red light at a constant frequency. In one embodiment, the light-emitting module may be placed on the first rim (1001) and / or the second rim (1002).
[0139] Referring to FIG. 10b, according to one embodiment, a wearable device (1099) may include a printed circuit board (PCB) (1090). The PCB (1090) may be included in at least one of a first temple (1004) or a second temple (1005). The PCB (1090) may include an interposer disposed between at least two sub-PCBs. On the PCB (1090), one or more hardware components included in the wearable device (1099) (e.g., hardware components illustrated by different blocks in FIG. 2) may be disposed. The wearable device (1099) may include a flexible PCB (FPCB) for interconnecting the hardware components.
[0140] According to one embodiment, a wearable device (1099) may include at least one of a gyroscope sensor, a gravity sensor, and / or an acceleration sensor for detecting the posture of the wearable device (1099) and / or the posture of a body part (e.g., head) of a user wearing the wearable device (1099). Each of the gravity sensor and the acceleration sensor may measure gravitational acceleration and / or acceleration based on designated three-dimensional axes (e.g., x-axis, y-axis, and z-axis) that are perpendicular to each other. The gyroscope sensor may measure the angular velocity of each of the designated three-dimensional axes (e.g., x-axis, y-axis, and z-axis). At least one of the gravity sensor, the acceleration sensor, and the gyroscope sensor may be referred to as an inertial measurement unit (IMU). According to one embodiment, the wearable device (1099) can identify a user's motion and / or gesture performed to execute or stop a specific function of the wearable device (1099) based on an IMU.
[0141] FIGS. 11a and 11b illustrate an example of the appearance of a wearable device (e.g., a wearable device (1099)). The wearable device (1099) of FIGS. 11a and 11b may be an example of the wearable device (1099) of FIG. 9. According to one embodiment, an example of the appearance of a first surface (1110) of the housing of the wearable device (1099) may be illustrated in FIG. 11a, and an example of the appearance of a second surface (1120) opposite to the first surface (1110) may be illustrated in FIG. 11b.
[0142] Referring to FIG. 11a, according to one embodiment, a first surface (1110) of a wearable device (1099) may have a shape that is attachable to a part of a user's body (e.g., the user's face). Although not illustrated, the wearable device (1099) may further include a strap for securing to a part of a user's body and / or one or more temples (e.g., a first temple (1004) and / or a second temple (1005) of FIG. 10a to FIG. 10b). A first display (1050-1) for outputting an image to the left eye among the user's two eyes, and a second display (1050-2) for outputting an image to the right eye among the two eyes may be disposed on the first surface (1110). The wearable device (1099) may further include rubber or silicone packing formed on the first surface (1110) to prevent interference by light different from light emitted from the first display (1050-1) and the second display (1050-2) (e.g., ambient light).
[0143] According to one embodiment, a wearable device (1099) may include cameras (1060-1) for photographing and / or tracking both eyes of a user adjacent to each of the first display (1050-1) and the second display (1050-2). The cameras (1060-1) may be referenced to the eye-tracking camera (1060-1) of FIG. 10b. According to one embodiment, a wearable device (1099) may include cameras (1060-5, 1060-6) for photographing and / or recognizing a user's face. The cameras (1060-5, 1060-6) may be referenced to FT cameras. The wearable device (1099) may control an avatar representing the user in a virtual space based on the motion of the user's face identified using the cameras (1060-5, 1060-6). For example, the wearable device (1099) can change the texture and / or shape of a part of an avatar (e.g., a part of an avatar representing a human face) by using information obtained by cameras (1060-5, 1060-6) (e.g., FT cameras) and representing the facial expression of a user wearing the wearable device (1099).
[0144] Referring to FIG. 11b, on a second surface (1120) opposite to the first surface (1110) of FIG. 11a, a camera (e.g., cameras (1060-7, 1060-8, 1060-9, 1060-10, 1060-11, 1060-12)), and / or a sensor (e.g., a depth sensor (1130)) may be disposed to obtain information related to the external environment of the wearable device (1099). For example, cameras (1060-7, 1060-8, 1060-9, 1060-10) may be disposed on the second surface (1120) to recognize external objects. The cameras (1060-7, 1060-8, 1060-9, 1060-10) can be referenced to the motion recognition cameras (1060-2, 1060-3) of FIG. 10b.
[0145] For example, using cameras (1060-11, 1060-12), the wearable device (1099) can acquire images and / or videos to be transmitted to each of the user's two eyes. Camera (1060-11) may be placed on the second surface (1120) of the wearable device (1099) to acquire an image to be displayed through a second display (1050-2) corresponding to the right eye among the two eyes. Camera (1060-12) may be placed on the second surface (1120) of the wearable device (1099) to acquire an image to be displayed through a first display (1050-1) corresponding to the left eye among the two eyes. Cameras (1060-11, 1060-12) may be referenced to the shooting camera (1060-4) of FIG. 10b.
[0146] According to one embodiment, the wearable device (1099) may include a depth sensor (1130) disposed on a second surface (1120) to identify the distance between the wearable device (1099) and an external object. Using the depth sensor (1130), the wearable device (1099) may obtain spatial information (e.g., a depth map) for at least a portion of the FoV of a user wearing the wearable device (1099). Although not illustrated, a microphone may be disposed on the second surface (1120) of the wearable device (1099) to obtain sound output from an external object. The number of microphones may be one or more, depending on the embodiment.
[0147] Hereinafter, with reference to FIG. 12, the hardware or software configuration of the wearable device (1099) will be described.
[0148] FIG. 12 illustrates an example of a block diagram of a wearable device (e.g., a wearable device (1099)). The wearable device (1099) of FIG. 12 may be an example of the electronic device (901) of FIG. 9 and the wearable device (1099) of FIG. 10a through FIG. 11b.
[0149] Referring to FIG. 12, a wearable device (1099) according to one embodiment may include a processor (1210), a memory (1215), a display (1050) (e.g., a first display (1050-1) and / or a second display (1050-2) of FIG. 10a, FIG. 10b, FIG. 11a, and FIG. 11b), and / or a sensor (1220). The processor (1210), memory (1215), display (1050) and / or sensor (1220) may be electrically and / or operationally connected to each other by an electronic component such as a communication bus (1202). In the present disclosure, the operational connection of the electronic components may include a direct connection established between the electronic components and / or an indirect connection established between the electronic components such that a first electronic component among the electronic components is controlled by a second electronic component among the electronic components. The type and / or number of electronic components included in the wearable device (1099) are not limited to those shown in FIG. 12. For example, the wearable device (1099) may include only some of the electronic components shown in FIG. 12.
[0150] A processor (1210) of a wearable device (1099) according to one embodiment may include a circuit (e.g., a processing circuit) for processing data based on one or more instructions. The circuit for processing data may include, for example, an arithmetic and logic unit (ALU), a field programmable gate array (FPGA), a central processing unit (CPU), and / or an application processor (AP). In one embodiment, the wearable device (1099) may include one or more processors. The processor (1210) may have a structure of a multi-core processor such as a dual core, a quad core, a hexa core, and / or an octa core. The multi-core processor structure of the processor (1210) may include a structure based on a plurality of core circuits (e.g., a big-little structure), distinguished by power consumption, clock, and / or computational power per unit time. In one embodiment comprising a processor (1210) having a multi-core processor structure, the operations and / or functions of the present disclosure may be performed individually or collectively by one or more cores included in the processor (1210).
[0151] A memory (1215) of a wearable device (1099) according to one embodiment may include electronic components for storing data and / or instructions that are input to or output from a processor (1210). The memory (1215) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, and embedded multi-media card (eMMC). In one embodiment, the memory (1215) may be referred to as storage.
[0152] In one embodiment, a display (1050) of a wearable device (1099) can output visualized information to a user of the wearable device (1099). A display (1050) arranged in front of the eyes of a user wearing the wearable device (1099) may be placed in at least a part of the housing of the wearable device (1099) (e.g., a first display (1050-1) and / or a second display (1050-2) of FIG. 10a, FIG. 10b, FIG. 11a, and FIG. 11b). For example, the display (1050) may be controlled by a processor (1210) including circuits such as a CPU, a GPU (graphic processing unit), and / or a DPU (display processing unit) to output visualized information to the user. The display (1050) may include a flexible display, a flat panel display (FPD), and / or electronic paper. The display (1050) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). Embodiments are not limited thereto, for example, if the wearable device (1099) includes a lens for transmitting external light (or ambient light), the display (1050) may include a projector (or projection assembly) for projecting light onto the lens. In one embodiment, the display (1050) may be referred to as a display panel and / or a display module. Pixels included in the display (1050) may be positioned toward either of the user's two eyes when worn by the user of the wearable device (1099).For example, the display (1050) may include display areas (or active areas) corresponding to each of the user's two eyes.
[0153] In one embodiment, a sensor (1220) of a wearable device (1099) may generate electrical information that can be processed by a processor (1210) and / or memory (1215) from non-electronic information associated with the wearable device (1099). For example, the sensor (1220) may include a global positioning system (GPS) sensor for detecting the geographic location of the wearable device (1099). In addition to the GPS method, the sensor (1220) may generate information indicating the geographic location of the wearable device (1099) based on a global navigation satellite system (GNSS), such as Galileo or Beidou (compass), for example. The above information may be stored in memory (1215), processed by a processor (1210), and / or transmitted to another electronic device distinct from the wearable device (1099) through a communication circuit.
[0154] According to one embodiment, within the memory (1215) of a wearable device (1099), one or more instructions (or commands) representing data to be processed by the processor (1210) of the wearable device (1099), calculations to be performed, and / or operations may be stored. A set of one or more instructions may be referred to as a program, firmware, operating system, process, routine, sub-routine, and / or software application (hereinafter, application). For example, the wearable device (1099), and / or processor (1210) may perform at least one of the operations of FIGS. 3 and FIGS. 5 when a set of a plurality of instructions distributed in the form of an operating system, firmware, driver, program, and / or software application is executed. In the following, the statement that a software application is installed in the wearable device (1099) means that one or more instructions provided in the form of a software application (or package) are stored in memory (1215), and that said one or more applications are stored in an executable format (e.g., a file having an extension specified by the operating system of the wearable device (1099)) by the processor (1210). For example, the application may include a program and / or library related to a service provided to a user.
[0155] Referring to FIG. 12, programs installed on a wearable device (1099) may be included in any one of different layers, including an application layer (1240), a framework layer (1250), and / or a hardware abstraction layer (HAL) (1280), based on the target. For example, within the hardware abstraction layer (1280), programs (e.g., modules, or drivers) designed to target the hardware of the wearable device (1099) (e.g., a display (1050), and / or a sensor (1220)) may be included. The framework layer (1250) may be referred to as an XR framework layer in that it includes one or more programs for providing XR (extended reality) services. For example, the layers illustrated in FIG. 12 are logically (or for convenience of explanation) separated, and this does not mean that the address space of memory (1215) is separated by said layers.
[0156] For example, within the framework layer (1250), programs designed to target at least one of the hardware abstraction layer (1280) and / or the application layer (1240) (e.g., a location tracker (1271), a spatial recognizer (1272), a gesture tracker (1273), an eye tracker (1274), and / or a face tracker (1275)) may be included. The programs included in the framework layer (1250) may provide an application programming interface (API) that is executable (or invokeable) based on other programs.
[0157] For example, within the application layer (1240), a program designed to target a user of the wearable device (1099) may be included. Examples of programs included in the application layer (1240) include an XR (extended reality) system UI (user interface) (1241) and / or an XR application (1242), but embodiments are not limited thereto. For example, programs included in the application layer (1240) (e.g., software applications) may call an API to cause the execution of a function supported by programs included in the framework layer (1250).
[0158] For example, the wearable device (1099) may display one or more visual objects on the display (1050) to perform interaction with the user based on the execution of the XR system UI (1241). A visual object may mean an object that can be placed on the screen for the transmission of information and / or interaction, such as text, images, icons, videos, buttons, checkboxes, radio buttons, text boxes, sliders, and / or tables. A visual object may be referred to as a visual guide, a virtual object, a visual element, a UI element, a view object, and / or a view element. The wearable device (1099) may provide the user with functions available in a virtual space based on the execution of the XR system UI (1241).
[0159] Referring to FIG. 12, a lightweight renderer (1243) and / or an XR plugin (1244) are depicted within the XR system UI (1241), but are not limited thereto. For example, based on the XR system UI (1241), the processor (1210) may execute a lightweight renderer (1243) and / or an XR plugin (1244) within a framework layer (1250).
[0160] For example, a wearable device (1099) may acquire resources (e.g., APIs, system processes, and / or libraries) used to define, create, and / or execute a rendering pipeline, which is permitted to be partially modified, based on the execution of a lightweight renderer (1243). The lightweight renderer (1243) may be referred to as a lightweight render pipeline in terms of defining a rendering pipeline, which is permitted to be partially modified. The lightweight renderer (1243) may include a renderer built prior to the execution of a software application (e.g., a pre-built renderer). For example, the wearable device (1099) may acquire resources (e.g., APIs, system processes, and / or libraries) used to define, create, and / or execute the entire rendering pipeline based on the execution of an XR plugin (1244). The XR plugin (1244) can be referred to as an open XR native client in terms of defining (or setting) the entire rendering pipeline.
[0161] For example, the wearable device (1099) may display a screen representing at least a portion of a virtual space on the display (1050) based on the execution of the XR application (1242). The XR plugin (1244-1) included in the XR application (1242) may include instructions that support functions similar to the XR plugin (1244) of the XR system UI (1241). Descriptions of the XR plugin (1244-1) that overlap with descriptions of the XR plugin (1244) may be omitted. The wearable device (1099) may trigger the execution of the virtual space manager (1251) based on the execution of the XR application (1242).
[0162] For example, the wearable device (1099) may display an image in virtual space on the display (1050) based on the execution of the application (1245). The application (1245) may be configured to output image information for displaying a two-dimensional image. The wearable device (1099) may trigger the execution of a virtual space manager (1251) based on the execution of the application (1245). The wearable device (1099) may generate dual image information to display the two-dimensional image in three-dimensional virtual space based on the execution of the application (1245). Here, the dual image information may include a first image information for the left eye and a second image information for the right eye, taking into account binocular parallax. To display the two-dimensional image in three-dimensional virtual space, the wearable device (1099) may generate the dual image information based on the image information for displaying the two-dimensional image.
[0163] According to one embodiment, the wearable device (1099) may provide a virtual space service based on the execution of a virtual space manager (1251). For example, the virtual space manager (1251) may include a platform for supporting the virtual space service. Based on the execution of the virtual space manager (1251), the wearable device (1099) may identify a virtual space formed based on the user's location indicated by data acquired through a sensor (1230), and may display at least a portion of the virtual space on a display (1050). The virtual space manager (1251) may be referred to as a composition presentation manager (CPM).
[0164] For example, the virtual space manager (1251) may include a runtime service (1252). For example, the runtime service (1252) may be referred to as an OpenXR runtime module (or OpenXR runtime program). The wearable device (1099) may execute at least one of a user pose prediction function, a frame timing function, and / or a spatial input function based on the execution of the runtime service (1252). For example, the wearable device (1099) may perform rendering for a virtual space service for the user based on the execution of the runtime service (1252). For example, a virtual space-related function executable by the application layer (1240) may be supported based on the execution of the runtime service (1252).
[0165] For example, the virtual space manager (1251) may include a pass-through manager (1253). The wearable device (1099) may display an image and / or video representing the real space acquired through an external camera superimposed on at least a portion of the screen while displaying a screen representing the virtual space on the display (1050) based on the execution of the pass-through manager (1253).
[0166] For example, the virtual space manager (1251) may include an input manager (1254). The wearable device (1099) may identify acquired data (e.g., sensor data) by executing one or more programs included within the recognition service layer (1270) based on the execution of the input manager (1254). The wearable device (1099) may identify user inputs associated with the wearable device (1099) using the acquired data. The user inputs may be associated with user motions (e.g., hand gestures), gaze, and / or speech identified by a sensor (1220) (e.g., an image sensor (1230) such as an external camera). The user inputs may be identified based on an external electronic device connected (or paired) via a communication circuit.
[0167] For example, the perception abstract layer (1260) can be used for data exchange between the virtual space manager (1251) and the perception service layer (1270). In terms of being used for data exchange between the virtual space manager (1251) and the perception service layer (1270), the perception abstract layer (1260) may be referred to as an interface. As an example, the perception abstract layer (1260) may be referred to as OpenPX. The perception abstract layer (1260) can be used for a perception client and a perception service.
[0168] According to one embodiment, the recognition service layer (1270) may include one or more programs for processing data obtained from the sensor (1220). The one or more programs may include at least one of a location tracker (1271), a spatial recognizer (1272), a gesture tracker (1273), and / or an eye tracker (1274). The type and / or number of the one or more programs included in the recognition service layer (1270) are not limited to those shown in FIG. 12.
[0169] For example, the wearable device (1099) can identify the posture of the wearable device (1099) using a sensor (1230) based on the execution of a position tracker (1271). The wearable device (1099) can identify the 6 degrees of freedom pose (6 dof pose) of the wearable device (1099) using data acquired using an external camera (e.g., image sensor (1221)) and / or an IMU (e.g., motion sensor (1222) including a gyroscope, accelerometer, and / or geomagnetic sensor) based on the execution of the position tracker (1271). The position tracker (1271) may be referred to as a head tracking (HeT) module (or head tracker, head tracking program).
[0170] For example, the wearable device (1099) may acquire information to provide a three-dimensional virtual space corresponding to the surrounding environment (e.g., external space) of the wearable device (1099) (or the user of the wearable device (1099)) based on the execution of the spatial recognizer (1272). The wearable device (1099) may reproduce the surrounding environment of the wearable device (1099) in three dimensions using data acquired using an external camera (e.g., image sensor (1221)) based on the execution of the spatial recognizer (1272). The wearable device (1099) may identify at least one of a plane, an incline, and a staircase based on the surrounding environment of the wearable device (1099) reproduced in three dimensions based on the execution of the spatial recognizer (1272). The space recognizer (1272) can be referred to as a scene understanding (SU) module (or scene understanding program).
[0171] For example, the wearable device (1099) can identify (or recognize) the pose and / or gesture of the user's hand of the wearable device (1099) based on the execution of the gesture tracker (1273). For example, the wearable device (1099) can identify the pose and / or gesture of the user's hand using data acquired from an external camera (e.g., image sensor (1221)) based on the execution of the gesture tracker (1273). For example, the wearable device (1099) can identify the pose and / or gesture of the user's hand based on data (or images) acquired using an external camera based on the execution of the gesture tracker (1273). The gesture tracker (1273) may be referred to as a hand tracking (HaT) module (or hand tracking program) and / or a gesture tracking module.
[0172] For example, the wearable device (1099) can identify (or track) the movement of the user's eyes of the wearable device (1099) based on the execution of the eye tracker (1274). For example, the wearable device (1099) can identify the movement of the user's eyes using data obtained from an eye-tracking camera (e.g., image sensor (1221)) based on the execution of the eye tracker (1274). The eye tracker (1274) may be referred to as an eye tracking (ET) module (or eye tracking program) and / or a gaze tracking module.
[0173] For example, the recognition service layer (1270) of the wearable device (1099) may further include a face tracker (1275) for tracking the user's face. For example, the wearable device (1099) may identify (or track) the movement of the user's face and / or the user's facial expression based on the execution of the face tracker (1275). The wearable device (1099) may estimate the user's facial expression based on the movement of the user's face based on the execution of the face tracker (1275). For example, the wearable device (1099) may identify the movement of the user's face and / or the user's facial expression based on data (e.g., images and / or videos) obtained using a camera (e.g., a camera facing at least a part of the user's face) based on the execution of the face tracker (1275).
[0174] Referring to FIG. 12, the renderer (1290) may include instructions for rendering images in a three-dimensional virtual space. A processor (1210) that executes the renderer (1290) may obtain at least one image to be displayed at least partially in a display area of the display (1050) in a software application. For example, the processor (1210) that executes the renderer (1290) may determine the location of the area where an application (e.g., XR application (1242), application (1245)) will be rendered. The processor (1210) that executes the renderer (1290) may generate an image of said application to be displayed on the display (1050). The renderer (1290) may synthesize images to generate a composite image to be displayed on the display (1050).
[0175] For example, a processor (1210) that executes a renderer (1290) can divide the display area of a display (1050) into a foveated portion (or may be referred to as a foveated area) and a peripheral portion (or may be referred to as a residual area) using a gaze position calculated using a position tracker (1271) and / or a gaze tracker (1274). For example, a processor (1210) that detects coordinate values of the gaze position can determine the portion of the display area containing said coordinate values as the foveated area. A DPU that executes a renderer (1290) can acquire at least one image corresponding to each of said foveated area and said residual area, having a size smaller than the size of the entire display area of the display (1050) or having a resolution less than the resolution of the display area.
[0176] A processor (1210) that executes a renderer (1290) can obtain or generate a composite image to be displayed on a display (1050) by synthesizing an image corresponding to a foveated area and an image corresponding to a surrounding area. For example, the processor (1210) can perform upscaling to enlarge the image corresponding to the surrounding area to the size of the entire display area of the display (1050). On the enlarged image, the processor (1210) can combine the image corresponding to the foveated area to generate a composite image to be displayed on the display (1050). Along the boundary line of the image corresponding to the foveated area, the processor (1210) can mix the enlarged image and the image corresponding to the foveated area by applying a visual effect such as blur.
[0177] FIG. 13 illustrates an example of a block diagram of an electronic device (e.g., electronic device (901), wearable device (1099)) for displaying an image in a virtual space. FIG. 13 describes an example in which multiple programs / instructions for displaying an image in a virtual space are executed. The multiple programs / instructions may all be executed on a single processor (e.g., AP) or may be executed by multiple processors (e.g., AP, GPU (graphic processing unit), NPU (neural processing unit)). The meaning of being able to be executed by multiple processors is that some programs / instructions may be executed by a first processor and other programs / instructions may be executed by a second processor different from the first processor.
[0178] Referring to FIG. 13, an electronic device (901) may execute a virtual space manager (1350) (e.g., the virtual space manager (1251) of FIG. 12, CPM) to render an image in a virtual space. For the virtual space manager (1350), at least some of the descriptions of the virtual space manager (1251) of FIG. 12 may be referenced. The virtual space manager (1350) may include a platform for supporting virtual space services. The virtual space manager (1350) may include a runtime service (1351) (e.g., OpenXR Runtime), a panel renderer (1352) (e.g., 2D Panel Render), and an XR composite unit (1353) (XR Compositor). Based on the execution of the runtime service (1351), the electronic device (901) may execute at least one of a user pose prediction function, a frame timing function, and / or a spatial input function. For the runtime service (1351), at least some of the descriptions of the runtime service (1252) of FIG. 12 may be referenced. The electronic device (901) may display at least one image (video) on a panel (e.g., a 2D panel) to enable the implementation of a virtual space through a display, based on the execution of the panel renderer (1352). For example, the electronic device (901) may display a rendering image corresponding to RGB information (1366) for the panel from the spatialization manager (1340) described later through a display (e.g., a display (1050)). The electronic device (901) may composite an image of a real area (hereinafter, a pass-through image) captured through a camera in virtual space with an image of a virtual area, based on the execution of the XR composite unit (1353) (XR Compositor). For example, the electronic device (901) can generate a composite image by merging the pass-through image and the virtual region image based on the execution of the XR synthesis unit (1353).The electronic device (901) can transmit the generated composite image to a display buffer so that the composite image is displayed. The electronic device (901) can identify a virtual space through a virtual space manager (1350) and display at least a portion of the virtual space on a display (1050). The virtual space manager (1350) may be referred to as CPM. The electronic device (901) can execute the virtual space manager (1350) to render an image corresponding to at least a portion of the virtual space.
[0179] According to one embodiment, an electronic device (901) may execute a spatialization manager (1340). The spatialization manager (1340) may perform processing for displaying an image in a three-dimensional virtual space. The electronic device (901) may perform preprocessing based on the execution of the spatialization manager (1340) so that an image can be rendered in a three-dimensional virtual space through a virtual space manager (1350). For example, the electronic device (901) may perform at least some of the functions of the renderer (1290) of FIG. 12 based on the execution of the spatialization manager (1340). The electronic device (901) may process image information provided by an application (e.g., an XR application (1310), an application providing a general 2D screen that is not XR (1320), an application providing a system UI (1330)) based on the execution of the spatialization manager (1340). A spatialization manager (1340) (e.g., Space Flinger) may include a system screen manager (1341) (e.g., System scene), an input manager (1342) (e.g., Input Routing), and a lightweight rendering engine (1343) (e.g., Impress Engine). The system screen manager (1341) may be executed to display a system UI (1330). System UI-related information (1364) may be transmitted to the system screen manager (1341) from a program (e.g., API) that provides the system UI (1330). System UI-related information (1364) may be obtained through a spatializer API and / or a Same-process private API. The spatialization manager (1340) may determine the layout (e.g., position, display order) of the system UI (1330) screen in three-dimensional space through pre-allocated resources.The system screen manager (1341) may transmit image information (1367) for rendering a screen of the system UI (1330) according to the layout to the virtual space manager (1350). The input manager (1342) may be configured to process user input (e.g., user input on a system screen or app screen). The lightweight rendering engine (1343) may be a renderer for image generation (e.g., lightweight renderer (1243)). For example, the lightweight rendering engine (1343) may be used to display the system UI (1330). According to one embodiment, the spatialization manager (1340) may include a lightweight rendering engine (1343) for rendering the system UI. According to one embodiment, if the lightweight rendering engine (1343) does not have sufficient resources to render an avatar used in an HMD, at least one external rendering engine may be used. At this time, to resolve compatibility issues with external rendering (e.g., 3rd party engine), an external rendering engine support module may be added inside the spatialization manager (1340).
[0180] According to one embodiment, the electronic device may execute an application. For example, in response to the execution of an XR application (1310) (e.g., an XR application (1242), a 3D game, an XR map, or other immersive application), the virtual space manager (1350) may be executed. The electronic device (901) may provide dual image information (1361) provided from the XR application (1310) to the virtual space manager (1350). To display images in three-dimensional space, the dual image information (1361) may include two image information that account for binocular parallax. For example, the dual image information (1361) may include a first image information for the user's left eye and a second image information for the user's right eye to render in three-dimensional virtual space. Hereinafter, the term dual image information is used in the present disclosure to refer to image information for displaying images for both eyes in three-dimensional space. In addition to the dual image information, the above dual image information may utilize binocular image information, dual image information, dual image data, dual image, binocular image data, stereoscopic image information, 3D image information, spatial image information, spatial image data, 2D-3D conversion data, dimension conversion image data, binocular parallax image data, and / or equivalent technical terms. The electronic device (901) can generate a composite image by merging image layers through a virtual space manager (1350). The electronic device (901) can transmit the generated composite image to a display buffer. The composite image can be displayed on the display (1050) of the electronic device (901).
[0181] According to one embodiment, the electronic device may execute at least one application among an XR application (1310) and other applications (1320) (e.g., a first application (1320-1), a second application (1320-2), ..., a Nth application (1320-N)). According to one embodiment, the application (1320) may be configured to output image information for displaying a two-dimensional image. In other words, the application (1320) may provide a two-dimensional image. For example, the application (1320) may be a video application, a schedule application, or an internet browser application. If, in response to the execution of the application (1320), image information (1362) provided from the application (1320) is provided to the virtual space manager (1350), the image information (1362) has only x-coordinates and y-coordinates within a two-dimensional plane, so it may be difficult to consider the sequential relationship between other applications centered on the user (i.e., distance from the user). The electronic device (901) may execute a spatialization manager (1340) to provide dual image information to a virtual space manager (1350), even when displaying an application (1320) that provides a general 2D screen. For example, based on the execution of the spatialization manager (1340), the electronic device (901) may receive application-related information (1363) from the first application (1320-1). For example, the application-related information (1363) may include image information representing a 2D image of the first application (1320-1) (e.g., information including RGB per pixel) and / or content information in the first application (1320-1) (e.g., characteristics of the content running in the first application, type of content). The application-related information (1363) may be obtained through a spatializer API.Based on the execution of the spatialization manager (1340), the electronic device (901) can identify information regarding the location of the area to be rendered and the size of the area to be rendered (hereinafter, location information). Based on the execution of the spatialization manager (1340), the electronic device (901) can generate dual image information (1365, e.g., RGBx2) that takes into account the user's binocular parallax through the image information and the location information. Based on the execution of the spatialization manager (1340), the electronic device (901) can provide the dual image information (1365) to the virtual space manager (1350). By converting a simple two-dimensional image into dual image information (1365), the problem caused by the image information (1362) being directly transmitted to the virtual space manager (1350) can be resolved. Additionally, as at least some of the functions for displaying images in virtual space are performed by the spatialization manager (1340) instead of the virtual space manager (1350), the burden on the virtual space manager (1350) may be reduced.
[0182] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.
[0183] An electronic device as described above (e.g., a head-wearable electronic device (100)) may include a memory (e.g., a memory (206)) for storing instructions. The electronic device may include one or more cameras (e.g., one or more first cameras (209)) available for acquiring images of the actual environment in front of the head-wearable electronic device. The electronic device may include a plurality of microphones (e.g., a plurality of microphones (204)). The electronic device may include at least one processor (e.g., at least one processor (207)). The instructions may cause the electronic device to acquire location information of a source object included in the actual environment using the images acquired through the one or more cameras when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, when the location information is acquired, the electronic device may be caused to acquire audio signals generated within the actual environment and characteristic information related to said audio signals through the plurality of microphones. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a target audio signal corresponding to said source object among said audio signals using said location information and said characteristic information.
[0184] According to one embodiment, the instructions may cause the electronic device to predict other characteristic information related to at least a portion of the audio signals to be acquired through the plurality of microphones, based on the location information, when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine that the at least a portion corresponds to the source object, based on the determination that the difference between the other characteristic information and the characteristic information is within a reference range, when executed individually or collectively by the at least one processor.
[0185] According to one embodiment, the characteristic information may represent at least one of a phase difference between audio signals obtained through the plurality of microphones, a difference in arrival time between the audio signals, and a difference in volume between the audio signals.
[0186] According to one embodiment, the head-wearable electronic device may further include a sensor. The position information is obtained through the sensor and may further include the distance between the head-wearable electronic device and the source object.
[0187] According to one embodiment, the one or more cameras may be one or more first cameras. The head-wearable electronic device may further include one or more second cameras arranged in relation to the eyes of the user of the head-wearable electronic device. The instructions may cause the electronic device to identify the user's gaze through the one or more second cameras when identifying a part of the body facing the head-wearable electronic device by performing object recognition on the images when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to add the body to a group of one or more candidate source objects based on acquiring the audio signals when identifying that the gaze is located on a visual object corresponding to the part of the body when executed individually or collectively by the at least one processor. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the target audio signal corresponding to the source object among the one or more candidate source objects by using the location information and the characteristic information.
[0188] According to one embodiment, the instructions may cause the electronic device to identify the state of the source object using the images when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to identify the audio signal obtained when identifying the target state using the location information and the characteristic information, based on a determination that the state is a target state for outputting audio, as the target audio signal corresponding to the source object.
[0189] According to one embodiment, the head-wearable electronic device may further include a display assembly. The instructions, when executed individually or collectively by the at least one processor, may cause the electronic device to display, through the display assembly, a visual representation for indicating the source object causing the other audio signals, based on identifying the target audio signal and acquiring other audio signals caused by the source object through the plurality of microphones.
[0190] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to determine the source object corresponding to the target audio signal among candidate source objects included in the actual environment, identified using the images by using the location information and the characteristic information.
[0191] A method performed by an electronic device (e.g., a head-wearable electronic device (100)) having one or more cameras (e.g., one or more first cameras (209)) and a plurality of microphones (e.g., a plurality of microphones (204)) available for acquiring images of a real environment in front of the head-wearable electronic device as described above may include an operation of acquiring location information for a source object included in the real environment using the images acquired through the one or more cameras. When acquiring the location information, the method may include an operation of acquiring audio signals generated within the real environment and characteristic information related to the audio signals through the plurality of microphones. The method may include an operation of identifying a target audio signal corresponding to the source object among the audio signals using the location information and the characteristic information.
[0192] According to one embodiment, the method may include an operation of predicting other characteristic information related to at least a portion of the audio signals to be acquired through the plurality of microphones based on the position information. The method may include an operation of determining that the at least a portion corresponds to the source object based on a determination that the difference between the other characteristic information and the characteristic information is within a reference range.
[0193] According to one embodiment, the characteristic information may represent at least one of a phase difference between audio signals obtained through the plurality of microphones, a difference in arrival time between the audio signals, and a difference in volume between the audio signals.
[0194] According to one embodiment, the head-wearable electronic device may further include a sensor. The position information is obtained through the sensor and may further include the distance between the head-wearable electronic device and the source object.
[0195] According to one embodiment, the one or more cameras may be one or more first cameras. The head-wearable electronic device may further include one or more second cameras arranged in relation to the eyes of the user of the head-wearable electronic device. The method may include the action of identifying the user's gaze through the one or more second cameras when identifying a part of the body facing the head-wearable electronic device by performing object recognition on the images. The method may include the action of adding the body to a group of one or more candidate source objects based on acquiring the audio signals when identifying that the gaze is located on a visual object corresponding to the part of the body. The method may include the action of identifying the target audio signal corresponding to the source object among the one or more candidate source objects using the location information and the characteristic information.
[0196] According to one embodiment, the head-wearable electronic device may further include a display assembly. The method may include, after identifying the target audio signal, an operation of displaying, through the display assembly, a visual representation for indicating the source object causing the other audio signals, based on acquiring other audio signals caused by the source object through the plurality of microphones.
[0197] According to one embodiment, the method may include an operation of determining the source object corresponding to the target audio signal among candidate source objects identified using the images and included in the actual environment, using the location information and the characteristic information.
[0198] In a computer-readable storage medium in which one or more programs are stored as described above, the one or more programs may include instructions that cause the electronic device (e.g., electronic device (100)) having one or more cameras (e.g., one or more first cameras (209)) and a plurality of microphones (e.g., a plurality of microphones (204)) available for acquiring images of a real environment in front of the head-wearable electronic device, to acquire location information of a source object included in the real environment using the images acquired through the one or more cameras. The one or more programs may include instructions that cause the electronic device to acquire audio signals generated within the real environment and characteristic information related to the audio signals through the plurality of microphones when the location information is acquired, when the electronic device is executed. The above one or more programs may include instructions that cause the electronic device to identify a target audio signal corresponding to the source object among the audio signals by using the location information and the characteristic information when executed by the electronic device.
[0199] According to one embodiment, the one or more programs may include instructions that cause the electronic device to predict other characteristic information related to at least a portion of the audio signals to be acquired through the plurality of microphones, based on the location information, when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine that the at least a portion corresponds to the source object, based on a determination that the difference between the other characteristic information and the characteristic information is within a reference range, when executed by the electronic device.
[0200] According to one embodiment, the characteristic information may represent at least one of a phase difference between audio signals obtained through the plurality of microphones, a difference in arrival time between the audio signals, and a difference in volume between the audio signals.
[0201] According to one embodiment, the head-wearable electronic device may further include a sensor. The position information is obtained through the sensor and may further include the distance between the head-wearable electronic device and the source object.
[0202] According to one embodiment, the one or more cameras may be one or more first cameras. The head-wearable electronic device may further include one or more second cameras arranged in relation to the eyes of the user of the head-wearable electronic device. The one or more programs may include instructions that cause the electronic device to identify the user's gaze through the one or more second cameras when identifying a part of the body facing the head-wearable electronic device by performing object recognition on the images when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to add the body to a group of one or more candidate source objects based on acquiring the audio signals when identifying that the gaze is located on a visual object corresponding to the part of the body when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to identify the target audio signal corresponding to the source object among the one or more candidate source objects by using the location information and the characteristic information when executed by the electronic device.
[0203] According to one embodiment, the one or more programs may include instructions that cause the electronic device to identify the state of the source object using the images when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to identify the audio signal obtained when identifying the target state using the location information and the characteristic information, based on a determination that the state is a target state for outputting audio when executed by the electronic device, as the target audio signal corresponding to the source object.
[0204] According to one embodiment, the head-wearable electronic device may further include a display assembly. The one or more programs may include instructions that cause the electronic device to display, through the display assembly, a visual representation for indicating the source object causing the other audio signals, based on identifying the target audio signal and then acquiring other audio signals caused by the source object through the plurality of microphones when executed by the electronic device.
[0205] According to one embodiment, the one or more programs may include instructions that cause the electronic device to determine the source object corresponding to the target audio signal among candidate source objects included in the actual environment, by using the location information and the characteristic information when executed by the electronic device.
[0206] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.
[0207] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0208] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0209] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a computer-executable program, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several combined hardware, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0210] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0211] Therefore, other implementations, other embodiments, and equivalents to the claims set forth below are also within the scope of the claims. According to one embodiment, the method according to the various embodiments disclosed herein may be provided as a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0212] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In a head-wearable electronic device, One or more cameras available for acquiring images of the actual environment in front of the head-wearable electronic device; Multiple microphones; Memory comprising one or more storage media for storing instructions; and It includes at least one processor comprising processing circuitry, and When the above instructions are executed individually or collectively by the at least one processor, Using the images obtained through the one or more cameras, location information for a source object included in the actual environment is obtained, and When the above location information is obtained, audio signals generated within the actual environment and characteristic information related to the audio signals are obtained through the plurality of microphones, and Using the above location information and the above characteristic information, to identify a target audio signal corresponding to the source object among the audio signals, Causing the above-mentioned head-wearable electronic device, Head-wearable electronic device.
2. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Based on the above location information, other characteristic information related to at least a portion of the audio signals to be acquired through the plurality of microphones is predicted, and Based on the determination that the difference between the above other characteristic information and the above characteristic information is within a reference range, to determine that at least some of the above correspond to the source object, Causing the above-mentioned head-wearable electronic device, Head-wearable electronic device.
3. In Claim 1, the characteristic information is, At least one of a phase difference between audio signals acquired through the plurality of microphones, a difference in arrival time between the audio signals, and a difference in volume between the audio signals. Head-wearable electronic device.
4. In claim 1, the head-wearable electronic device is, Includes additional sensors, and The above location information is, Acquired through the sensor above, further including the distance between the head-wearable electronic device and the source object, Head-wearable electronic device.
5. In claim 1, the one or more cameras are, One or more first cameras, and The above head-wearable electronic device is, The above head-wearable electronic device further includes one or more second cameras arranged in relation to the user's eyes, and When the above instructions are executed individually or collectively by the at least one processor, When identifying a part of the body facing the head-wearable electronic device by performing object recognition on the above images, the user's gaze is identified through the one or more second cameras, and When it is identified that the gaze is positioned on a visual object corresponding to the part of the body, based on acquiring the audio signals, the body is added to a group of one or more candidate source objects, and Using the above location information and the above characteristic information, to identify the target audio signal corresponding to the source object among the one or more candidate source objects, Causing the above-mentioned head-wearable electronic device, Head-wearable electronic device.
6. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Using the above images, identify the state of the source object, and Based on the determination that the above state is a target state for outputting audio, using the location information and the characteristic information, the audio signal obtained when identifying the target state is identified as the target audio signal corresponding to the source object. Causing the above-mentioned head-wearable electronic device, Head-wearable electronic device.
7. In claim 1, the head-wearable electronic device is, It further includes a display assembly, When the above instructions are executed individually or collectively by the at least one processor, After identifying the target audio signal, based on acquiring other audio signals caused by the source object through the plurality of microphones, a visual representation for indicating the source object causing the other audio signals is to be displayed through the display assembly. Causing the above-mentioned head-wearable electronic device, Head-wearable electronic device.
8. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, Using the above location information and the above characteristic information, to determine the source object corresponding to the target audio signal among the candidate source objects identified using the above images and included in the actual environment, Causing the above-mentioned head-wearable electronic device, Head-wearable electronic device.
9. In a non-transient computer-readable storage medium storing one or more programs, said one or more programs, said one or more programs are, When executed by a head-wearable electronic device having one or more cameras and a plurality of microphones available to acquire images of the actual environment in front of the head-wearable electronic device, Using the images obtained through the one or more cameras, location information for a source object included in the actual environment is obtained, and When the above location information is obtained, audio signals generated within the actual environment and characteristic information related to the audio signals are obtained through the plurality of microphones, and Using the above location information and the above characteristic information, to identify a target audio signal corresponding to the source object among the audio signals, Instructions for causing the above-mentioned head-wearable electronic device, Non-transient computer-readable storage media.
10. In Claim 9, When the above one or more programs are executed by the head-wearable electronic device, Based on the above location information, other characteristic information related to at least a portion of the audio signals to be acquired through the plurality of microphones is predicted, and Based on the determination that the difference between the above other characteristic information and the above characteristic information is within a reference range, to determine that at least some of the above correspond to the source object, Instructions for causing the above-mentioned head-wearable electronic device, Non-transient computer-readable storage media.
11. In claim 9, the characteristic information is, At least one of a phase difference between audio signals acquired through the plurality of microphones, a difference in arrival time between the audio signals, and a difference in volume between the audio signals. Non-transient computer-readable storage media.
12. In claim 9, the head-wearable electronic device is, Includes additional sensors, and The above location information is, Acquired through the sensor above, further including the distance between the head-wearable electronic device and the source object, Non-transient computer-readable storage media.
13. In claim 9, the one or more cameras are, One or more first cameras, and The above head-wearable electronic device is, The above head-wearable electronic device further includes one or more second cameras arranged in relation to the user's eyes, and When the above one or more programs are executed by the head-wearable electronic device, When identifying a part of the body facing the head-wearable electronic device by performing object recognition on the above images, the user's gaze is identified through the one or more second cameras, and When it is identified that the gaze is positioned on a visual object corresponding to the part of the body, based on acquiring the audio signals, the body is added to a group of one or more candidate source objects, and Using the above location information and the above characteristic information, to identify the target audio signal corresponding to the source object among the one or more candidate source objects, Instructions for causing the above-mentioned head-wearable electronic device, Non-transient computer-readable storage media.
14. In Claim 9, When the above one or more programs are executed by the head-wearable electronic device, Using the above images, identify the state of the source object, and Based on the determination that the above state is a target state for outputting audio, using the location information and the characteristic information, the audio signal obtained when identifying the target state is identified as the target audio signal corresponding to the source object. Instructions for causing the above-mentioned head-wearable electronic device, Non-transient computer-readable storage media.
15. In claim 9, the head-wearable electronic device is, It further includes a display assembly, When the above one or more programs are executed by the head-wearable electronic device, After identifying the target audio signal, based on acquiring other audio signals caused by the source object through the plurality of microphones, a visual representation for indicating the source object causing the other audio signals is to be displayed through the display assembly. Instructions for causing the above-mentioned head-wearable electronic device, Non-transient computer-readable storage media.