Head-wearable electronic device, method, and non-transitory computer-readable storage medium for determining user preference information

WO2026168822A1PCT designated stage Publication Date: 2026-08-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-08-13

Smart Images

  • Figure KR2026001181_13082026_PF_FP_ABST
    Figure KR2026001181_13082026_PF_FP_ABST
Patent Text Reader

Abstract

This head-wearable electronic device comprises: at least one processor including processing circuitry; one or more cameras; one or more sensors; and a memory storing one or more programs configured to be executed individually or collectively by the at least one processor, and including one or more storage media, wherein the one or more programs may include instructions that cause the head-wearable electronic device to: drive the one or more cameras; while the one or more cameras are being driven, acquire sensing data from the one or more sensors; determine, on the basis of the sensing data, whether an object in a field of view (FOV) of the one or more cameras is an object of interest; add, on the basis of determining that the object is the object of interest, data regarding the object to a database; and determine, on the basis of the database, user preference information to be provided to a model trained with a prompt to refine a response to the prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Head-worn electronic device, method, and non-transient computer-readable storage medium for determining user preference information

[0001] The present disclosure relates to a head-worn electronic device, a method, and a non-transient computer-readable storage medium for determining user preference information.

[0002] To provide an enhanced user experience, electronic devices are being developed that provide augmented reality (AR) services by displaying computer-generated information in conjunction with external objects within the real world. The electronic device may be a head-mounted electronic device that can be worn by a user. For example, the electronic device may be AR glasses and / or a head-mounted device (HMD).

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] A head-wearing electronic device is described. The head-wearing electronic device may include at least one processor comprising a processing circuit, one or more cameras, one or more sensors, and memory comprising one or more storage media for storing one or more programs configured to be executed individually or collectively by the at least one processor. The one or more programs may include instructions that cause the head-wearing electronic device to drive the one or more cameras. The one or more programs may include instructions that cause the head-wearing electronic device to acquire sensing data from the one or more sensors while driving the one or more cameras. The one or more programs may include instructions that cause the head-wearing electronic device to determine whether an object within the field of view (FOV) of the one or more cameras is an object of interest based on the sensing data. The one or more programs above may include instructions that cause the head-worn electronic device to add data regarding the object into a database based on determining that the object is the object of interest. The one or more programs above may include instructions that cause the head-worn electronic device to determine user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt based on the database.

[0005] A method is described. The method may be performed within a head-worn electronic device comprising one or more cameras and one or more sensors. The method may include an operation of driving the one or more cameras. The method may include an operation of acquiring sensing data from the one or more sensors while driving the one or more cameras. The method may include an operation of determining, based on the sensing data, whether an object within the field of view (FOV) of the one or more cameras is an object of interest. The method may include an operation of adding data regarding the object to a database based on determining that the object is an object of interest. The method may include an operation of determining, based on the database, user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt.

[0006] A non-transient computer-readable storage medium is described. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the head-wearing electronic device to drive the one or more cameras when executed by the head-wearing electronic device having one or more cameras and one or more sensors. The one or more programs may include instructions that cause the head-wearing electronic device to acquire sensing data from the one or more sensors while driving the one or more cameras when executed by the head-wearing electronic device. Based on the sensing data, the head-wearing electronic device may include instructions that cause the head-wearing electronic device to determine whether an object within the field of view (FOV) of the one or more cameras is an object of interest. The one or more programs described above may include instructions that cause the head-wearing electronic device to add data regarding the object into a database based on determining that the object is the object of interest when executed by the head-wearing electronic device. The one or more programs described above may include instructions that cause the head-wearing electronic device to determine, based on the database, user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt when executed by the head-wearing electronic device.

[0007] Figure 1 illustrates an example of displaying information for an object using a trained model.

[0008] Figure 2 is a simplified block diagram of an exemplary head-worn electronic device.

[0009] FIG. 3 is a flowchart illustrating exemplary operations of a head-worn electronic device for adding data about an object into a database.

[0010] Figure 4 illustrates an example of determining whether an object is an object of interest based on sensing data.

[0011] Figure 5 illustrates a vector database containing embedding vectors.

[0012] FIG. 6 is a flowchart illustrating exemplary operations of a head-worn electronic device for determining user preference information.

[0013] Figure 7 illustrates a vector database and other vector databases.

[0014] FIG. 8 is a flowchart illustrating exemplary operations of a head-worn electronic device for displaying information for other objects.

[0015] Figure 9 illustrates an example of displaying information for other objects.

[0016] FIG. 10 is a flowchart illustrating exemplary operations of a head-worn electronic device for deleting data regarding an object in a database.

[0017] FIG. 11 is a block diagram of an electronic device in a network environment according to various embodiments.

[0018] FIG. 12 illustrates an example of a generative artificial intelligence system according to one embodiment.

[0019] FIG. 13a shows an example of a perspective view of a wearable device.

[0020] FIG. 13b shows an example of one or more hardware components placed within a wearable device.

[0021] FIGS. 14a and 14b show examples of the appearance of a wearable device.

[0022] Figure 15 shows an example of a block diagram of a wearable device.

[0023] Figure 16 shows an example of a block diagram of an electronic device for displaying an image in virtual space.

[0024] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0025] Figure 1 illustrates an example of displaying information for an object using a trained model.

[0026] Referring to FIG. 1, the head-wearing electronic device (100) may include a head-mounted display (HMD) that can be worn on the head of a user (110). The electronic device (100) may be described as a head-mounted display (HMD) device, a headgear electronic device, a glasses-type (or goggle-type) electronic device, a video see-through (VST) or visible see-through (VST) device, an extended reality (XR) device, a virtual reality (VR) device, and / or an augmented reality (AR) device.

[0027] A head-wearing electronic device (100) may include one or more cameras (e.g., one or more cameras (240) of FIG. 2) and a display assembly (e.g., a display assembly (230) of FIG. 2). In state (105), the head-wearing electronic device (100) may acquire an image representing an external space (or external environment) within the field of view (FOV) of one or more cameras through one or more cameras. For example, the image may include a representation of an object (e.g., object (120), object (125-1), and / or object (125-2)) within the field of view (FOV) of one or more cameras. The head-wearing electronic device (100) may display the image through the display assembly (230). The head-worn electronic device (100) can allow the user (110) to view the external space (or external environment) within the FOV of one or more cameras by displaying the image.

[0028] For example, within an external space (or external environment), an object (120) may be located relatively far from the head-worn electronic device (100) compared to objects (125-1) and (125-2). For example, as the object (120) is located relatively far from the head-worn electronic device (100) compared to objects (125-1) and (125-2), the representation of the object (120) within the image may be displayed relatively smaller than the representations of objects (125-1) and (125-2). For example, as the representation of the object (120) within the image is displayed relatively smaller than the representations of objects (125-1) and (125-2), a user (110) attempting to view the object (125-1) may experience discomfort. For example, a trained model may be used to alleviate the user's (110) discomfort.

[0029] A head-worn electronic device (100) can generate a prompt from a speech (130) of a user (110). For example, the speech (130) of the user (110) can be identified through a microphone of the head-worn electronic device (100) (e.g., microphone (260) of FIG. 2). For example, the prompt may include text requesting information about an external environment (or external space) (e.g., “Tell me what the situation is”). The head-worn electronic device (100) can provide the prompt and the image representing the external environment (or external space) to a trained model. The head-worn electronic device (100) can obtain information (135) for the external environment (or external space) generated by the trained model using the prompt and the image. The head-worn electronic device (100) can provide a response to the user's (110) speech (130) by displaying information (135) for the external environment (or external space) through a display assembly. For example, the information (135) for the external environment (or external space) may include information for an object (125-1) and / or an object (125-2) (e.g., text such as “A and B are eating and talking”), and may not include information for an object (120) (e.g., an image representing an enlarged object (120)). For example, since only a prompt is provided to the trained model, the information (135) generated using the trained model may not include information for the object (120) that the user wants (or intends). For example, since the user's gaze (140) is directed toward the object (120) within the state (105), the user may want information for the object (120). Since the information (135) for the external environment (or external space) does not include information for the object (120), a user who wants (or intends) information for the object (120) may feel inconvenienced.For example, a solution may be required to resolve the inconvenience to the user caused by not displaying information for the object (120).

[0030] To resolve this inconvenience, the head-wearing electronic device (100) may use user preference information provided to a trained model along with the prompt to refine the response regarding the prompt. For example, the head-wearing electronic device (100) may determine user preference information using data regarding an object (120). For example, the head-wearing electronic device (100) may obtain information for an object by providing a prompt regarding the object (120) along with user preference information to a trained model.

[0031] A head-worn electronic device (100) may perform operations exemplified within the description of FIGS. 3 through 10 to determine user preference information. The head-worn electronic device (100) may include components for performing said operations. said components may be exemplified within the description of FIG. 2.

[0032] Figure 2 is a simplified block diagram of an exemplary head-worn electronic device.

[0033] Referring to FIG. 2, according to one embodiment, a head-wearing electronic device (200) may be described as a head-mount display (HMD) device, headgear electronic device, glasses-type (or goggle-type) electronic device, video see-through (VST) device, extended reality (XR) device, virtual reality (VR) device, and / or augmented reality (AR) device that can be worn on a user's head. The head-wearing electronic device (200) may be worn on a user's head. An example of the structure of the electronic device (200) is described with reference to FIG. 13a, FIG. 13b, FIG. 14a and / or FIG. 14b. The head-wearing electronic device (200) may include at least a part of the head-wearing electronic device (100) of FIG. 1 or correspond to at least a part of the electronic device (100) of FIG. 1. The head-wearing electronic device (200) may include at least a part of the electronic device (1101) of FIG. 11 or correspond to at least a part of the electronic device (1101) of FIG. 11. The head-wearing electronic device (200) may include at least one processor (210) (e.g., processor (1120) of FIG. 11), memory (220) (e.g., memory (1130) of FIG. 11), display assembly (230) (e.g., display module (1160) of FIG. 11), one or more cameras (240) (e.g., camera module (1180) of FIG. 11), one or more sensors (250) (e.g., sensor module (1176) of FIG. 11), and / or a microphone (260) (e.g., input module (1150) of FIG. 11).

[0034] At least one processor (210) may include a processing circuit. At least one processor (210) may include a central processing unit (e.g., including a processing circuit). At least one processor (210) may include a graphic processing unit (e.g., including a processing circuit) and a neural processing unit (e.g., including a processing circuit). For example, at least one processor (210) may be configured to control a memory (220), a display assembly (230), one or more cameras (240), one or more sensors (250), and / or a microphone (260). At least one processor (210) may be configured to execute instructions stored in memory (220) individually or collectively to cause an electronic device (200) (or electronic device (100)) to perform at least some of the operations illustrated in the description of FIG. 1. At least one processor (210) may be configured to execute instructions stored in memory (220) individually or collectively to cause the electronic device (200) to perform at least some of the operations to be illustrated in the description of FIGS. 3 through 10.

[0035] For example, the term “processor” as used herein, including in the claims, may include various processing circuits comprising at least one processor, and one or more of said at least one processor may be configured to perform the various functions described below in a distributed manner, individually and / or collectively. As used below, where “processor,” “at least one processor,” and “one or more processors” are described as being configured to perform various functions, these terms encompass, for example, but not limited to, situations where one processor performs some of the cited functions and another processor(s) perform other parts of the cited functions, and also situations where one processor can perform all of the cited functions. Additionally, said at least one processor may include a combination of processors that perform the enumerated / disclosed various functions, for example, in a distributed manner. At least one processor may execute program instructions to achieve or perform the various functions.

[0036] The memory (220) may include one or more storage media. The memory (220) may store various data used by at least one component of the electronic device (200) (e.g., at least one processor (210), a display assembly (230), one or more cameras (240), one or more sensors (250), and / or a microphone (260)). For example, the data may include input data or output data for software and related commands. The memory (220) may include volatile memory or non-volatile memory.

[0037] The display assembly (230) may be configured to visualize information (or signals) provided from at least one processor (210). The display assembly (230) may be positioned toward the eyes of a user wearing the electronic device (200). The display assembly (230) may be configured to display at least a portion of images acquired through one or more cameras (240). The display assembly (230) may be configured to display information. The display assembly (230) may include at least one display.

[0038] One or more cameras (240) may include one or more light sensors (e.g., a charged coupled device (CCD) sensor and / or a complementary metal oxide semiconductor (CMOS) sensor) that generate an electrical signal representing the color and / or brightness of light. For example, one or more cameras (240) may be described as image sensors. For example, one or more cameras (240) may be available to acquire images representing the space (or external environment) outside the electronic device (200). For example, one or more cameras (240) may have a field of view (FOV). For example, at least a portion of one or more cameras (240) may have a field of view (FOV) corresponding to the field of view (FOV) of the user's eye. For example, the FOV of a portion of one or more cameras (240) may be different from the FOV of another portion of one or more cameras (240).

[0039] One or more sensors (250) may be configured to identify the posture (or orientation) of the head-worn electronic device (200). One or more sensors (250) may identify changes in speed caused by the movement of the head-worn electronic device (200). For example, one or more sensors (250) may obtain depth information of the object by using the difference between the light emitted by light-emitting elements and the light received by light-receiving elements. For example, one or more sensors (250) may include an acceleration sensor or an accelerometer. One or more sensors (250) may include a gyro sensor. For example, one or more sensors (250) may include one or more cameras arranged toward the user's eyes. For example, one or more sensors (250) may be used to identify the user's gaze. For example, one or more sensors (250) can be coupled with at least one processor (210) operably or operatively.

[0040] A microphone (260) may be used to convert sound into an electrical signal. For example, the microphone (260) may be configured to acquire an audio signal generated from outside the head-wearing electronic device (200). For example, the microphone (260) may include a digital microphone, an electronic condenser microphone (ECM), or a micro electro mechanical system (MEMS). However, it is not limited thereto.

[0041] The head-wearing electronic device (200) illustrated in the description of FIG. 2 may perform at least some of the operations illustrated in the descriptions of FIG. 3 through 10. The operations illustrated in the descriptions of FIG. 3 through 10 may be caused by (or within) the head-wearing electronic device (200) under the control of at least one processor (210).

[0042] FIG. 3 is a flowchart illustrating exemplary operations of a head-worn electronic device for adding data about an object into a database.

[0043] Referring to FIG. 3, according to one embodiment, in operation 300, at least one processor (210) can drive one or more cameras (240). While driving one or more cameras (240), the at least one processor (210) can acquire sensing data from one or more sensors (250). For example, the sensing data may include data for the posture (or orientation) of the head-worn electronic device (200), data for the user's gaze, and / or data for the user's face (or facial expression). For example, the sensing data may be referred to as perception data. For example, the perception data may include data for proprioception, data for vestibular sense, data for external light, and / or temperature data. However, it is not limited thereto.

[0044] According to one embodiment, at least one processor (210) can track a user's hand using one or more cameras (240) while driving one or more cameras (240). For example, at least one processor (210) can obtain data for the user's hand by tracking the user's hand. For example, at least one processor (210) can obtain an image representing the external environment within the FOV of one or more cameras (240) using one or more cameras (240) while driving one or more cameras (240). At least one processor (210) can identify the external environment represented by the image using a trained model (e.g., a scene understanding model). By identifying the external environment, at least one processor (210) can obtain data for the external environment. For example, the cognitive data may include data for the user's hand and / or data for the external environment. For example, at least one processor (210) can obtain an audio signal originating from outside the head-wearing electronic device (200) while driving one or more cameras (240). For example, the cognitive data may include the audio signal. For example, at least one processor (210) can obtain text for the audio signal by performing speech-to-text (STT) on the audio signal originating from outside the head-wearing electronic device (200). For example, the cognitive data may include text for the audio signal obtained by performing STT, but is not limited thereto.

[0045] According to one embodiment, at least one processor (210) can identify an object within the FOV of one or more cameras (240) using one or more cameras (240). For example, the object may be identified as it moves from outside the FOV of one or more cameras (240) into the FOV of one or more cameras (240), or as the direction in which the head-worn electronic device (200) is facing changes. For example, the object may be located in an external space (or external environment) within the FOV of one or more cameras (240).

[0046] According to one embodiment, in operation 310, at least one processor (210) can determine whether an object within the FOV of one or more cameras (240) is an object of interest based on sensing data (or perception data). For example, an object of interest may be described as an object perceived by a user or an object reacted to by a user. Determining whether an object within the FOV of one or more cameras (240) is an object of interest is illustrated in the description of FIG. 4.

[0047] Figure 4 illustrates an example of determining whether an object is an object of interest based on sensing data.

[0048] Referring to FIG. 4, according to one embodiment, a state (400) may be described as a state in which an object (415) in an external space (or external environment) is located within the FOV of one or more cameras (240). In the state (400), at least one processor (210) can identify an object (415) within the FOV of one or more cameras (240) using one or more cameras (240). At least one processor (210) can determine whether the object (415) within the FOV of one or more cameras (240) is an object of interest based on sensing data (or perception data) acquired while driving one or more cameras (240).

[0049] According to one embodiment, at least one processor (210) may acquire sensing data from one or more sensors (250) while identifying an object (415) within the FOV of one or more cameras (240). For example, at least one processor (210) may use the sensing data to identify that the head-worn electronic device (200) has a posture (or orientation) toward the object (415). For example, at least one processor (210) may determine that the object (415) is an object of interest (or the object (415) is an object of interest) based on identifying that the head-worn electronic device (200) has a posture (or orientation) toward the object (415). For example, at least one processor (210) may use the sensing data to identify that the neck of the user (405) is stretched toward the object (415). For example, at least one processor (210) may determine that the object (415) is an object of interest based on identifying that the user (405)'s neck is stretched toward the object (415). For example, at least one processor (210) may use sensing data to identify that the user (405)'s gaze (420) is directed toward the object (415). For example, at least one processor (210) may determine that the object (415) is an object of interest based on identifying that the user (405)'s gaze (420) is directed toward the object (415). However, it is not limited thereto.

[0050] According to one embodiment, at least one processor (210) may acquire an audio signal (425) generated from an object (415) and / or an audio signal (430) generated from a user (405) through a microphone (260) while identifying an object (415) within the FOV of one or more cameras (240). For example, at least one processor (210) may determine that the object (415) is an object of interest based on the audio signal (425) and / or audio signal (430) acquired through the microphone (260).

[0051] According to one embodiment, at least one processor (210) can track the hand of a user (405) through one or more cameras (240) while identifying an object (415) within the FOV of one or more cameras (240). For example, by tracking the hand of the user (405), the at least one processor (210) can identify that the user's (405) hand (or fingertip) points to the object (415) or that the user's (405) hand (or fingertip) is directed toward the object (415). For example, the at least one processor (210) can determine that the object (415) is an object of interest based on identifying that the user's (405) hand (or fingertip) points to the object (415) or that the user's (405) hand (or fingertip) is directed toward the object (415).

[0052] According to one embodiment, cognitive data (e.g., data for proprioception, data for vestibular sensation, data for external light, and / or temperature data) may be further used to determine whether the object (415) is an object of interest. However, it is not limited thereto.

[0053] Referring again to FIG. 3, according to one embodiment, in operation 320, at least one processor (210) can identify whether a database contains a set of data related to said object based on determining that an object within the FOV of one or more cameras (240) is an object of interest. For example, the database may be created by accumulating data regarding a plurality of objects. For example, data regarding a plurality of objects may be referred to as context data. For example, data regarding a plurality of objects may include embedding vectors. For example, embedding vectors may be defined as values ​​of real vectors representing text, images, and / or audio. For example, the database may include a vector database (or vector store) containing data regarding a plurality of objects. For example, the vector database (or vector store) may be defined as a database configured to efficiently store embedding vectors. For example, the database may include a plurality of sets of data. For example, data regarding multiple objects may be divided into multiple sets of data in a database based on the similarity between the data regarding multiple objects. For example, data regarding objects with relatively high similarity may be divided into the same set, and data regarding objects with relatively low similarity may be divided into different sets. For example, a set of data related to an object within the FOV of one or more cameras (240) may be defined as a set of data including data regarding objects with high similarity to said object.

[0054] According to one embodiment, in operation 330, at least one processor (210) may add data regarding the object to the set of data regarding the object included in the database based on identifying that the database includes a set of data related to the object. For example, the data regarding the object may be referred to as context data. For example, at least one processor (210) may acquire data regarding the object based on sensing data (or cognitive data). For example, the data regarding the object may include an embedding vector. For example, at least one processor (210) may extract features of the object by combining (or concatenating) the sensing data (or cognitive data). For example, the embedding vector may be acquired by providing the features of the object to a trained model. For example, the data regarding the object may include a spectrogram of an audio signal (e.g., an audio signal generated by the object and / or an audio signal generated by the user) acquired through a microphone (260). For example, the spectrogram of an audio signal can be obtained by providing the audio signal to a CNN (convolutional neural network) model. However, it is not limited to this.

[0055] According to one embodiment, data regarding an object added to a set of data related to the object may be stored in a database in conjunction with data regarding objects within the set of data related to the object. For example, by adding data regarding an object to a set of data related to the object included in the database, data regarding the object may be accumulated in the database.

[0056] According to one embodiment, in operation 340, at least one processor (210) may create a set of data related to the object in the database based on identifying that the database does not contain a set of data related to the object. For example, a set of data related to the object may be newly created to add data related to the object to the database.

[0057] According to one embodiment, in operation 350, at least one processor (210) may add data regarding the object to the set of generated data. For example, the data regarding the object added to the set of generated data may be stored independently of other data in the database. For example, by adding data regarding the object to the set of data related to the object included in the database, data regarding the object may be accumulated in the database. Adding data related to the object, which is an embedding vector, to the vector database is exemplified in the description of FIG. 5.

[0058] Figure 5 illustrates a vector database containing embedding vectors.

[0059] Referring to FIG. 5, according to one embodiment, the database is a vector database (or vector store) (500), and the data of the database may be embedding vectors (505). FIG. 5 illustrates the database as a vector database (500) and the data as embedding vectors (505) for convenience, but this is exemplary and is not limited thereto. The distance between the embedding vectors (505) within the vector database (500) may indicate the similarity between the objects indicated by the embedding vectors (505). For example, the distance between the embedding vectors (505) within the vector database (500) may be inversely proportional to the similarity between the objects indicated by the embedding vectors (505).

[0060] According to one embodiment, embedding vectors (505) within a vector database (500) may be represented in different colors depending on the set of embedding vectors. For example, the colors of the embedding vectors (505) may indicate perspective information. For example, embedding vectors (505-1) having a first color may be included in a first set of embedding vectors, embedding vectors (505-2) having a second color may be included in a second set of embedding vectors, embedding vectors (505-3) having a third color may be included in a third set of embedding vectors, and embedding vectors (505-4) having a fourth color may be included in a fourth set of embedding vectors. For example, as embedding vectors of the same color are accumulated within the vector database (500), perspective information may be generated (or formed). For example, embedding vectors (505) may be distributed among multiple sets of embedding vectors in a vector database based on the similarity between objects indicated by the embedding vectors (505). For example, embedding vectors for objects with high similarity may be distributed within the same set of embedding vectors, and embedding vectors for objects with low similarity may be distributed within different sets of embedding vectors. For example, embedding vectors distributed within a specific set of embedding vectors may have the same orientation. For example, as embedding vectors having the same orientation are accumulated in the vector database (500), perspective information may be generated (or formed).

[0061] According to one embodiment, at least one processor (210) can identify whether a vector database contains a set of embedding vectors associated with an object within the FOV of one or more cameras (240). For example, the object may have a relatively high degree of similarity to objects indicated by embedding vectors (505-2) in a second set of embedding vectors. For example, as the degree of similarity between the object and objects indicated by embedding vectors (505-2) in a second set of embedding vectors is relatively high, the object may be associated with a second set of embedding vectors. For example, at least one processor (210) can identify a second set of embedding vectors associated with the object among the sets of embedding vectors included in the vector database (500). For example, at least one processor (210) may add an embedding vector (510-1) regarding the object to a second set of embedding vectors regarding the object based on identifying a second set of embedding vectors related to the object. For example, the embedding vector (510-1) regarding the object added to the second set of embedding vectors may be stored in the vector database (500) in association with the embedding vectors (505-2) in the second set of embedding vectors. For example, the embedding vector (510-1) regarding the object may be added in the vector database (500) at a location adjacent to the embedding vectors (505-2) in the second set of embedding vectors. For example, in the vector database (500), the embedding vector (510-1) for the object may have a second color corresponding to the color of the embedding vectors (505-2) in the second set of embedding vectors.

[0062] According to one embodiment, the vector database (500) may not include embedding vectors for objects that have a relatively high similarity to the object. For example, as the vector database (500) does not include embedding vectors for objects that have a relatively high similarity to the object, the sets of embedding vectors included in the vector database (500) may not include a set of embedding vectors associated with the object. At least one processor (210) may generate a fifth set of embedding vectors associated with the object in the vector database (500) based on identifying that the vector database does not include a set of embedding vectors associated with the object. At least one processor (210) may add an embedding vector (510-2) for the object to the fifth set of embedding vectors. For example, the embedding vector (510-2) for the object added to the fifth set of embedding vectors may be separated from the embedding vectors (505) in the vector database (500). For example, the embedding vector (510-2) for the above object may have a fifth color different from the color of the embedding vectors (505) in the vector database (500).

[0063] According to one embodiment, at least one processor (210) can determine user preference information based on a database to which data regarding objects has been added. For example, as user preference information accumulates, the accumulated user preference information may be referred to as perspective information. The user preference information may be provided to a model trained with prompts to refine responses regarding prompts. Determining user preference information is exemplified in the description of FIG. 6.

[0064] FIG. 6 is a flowchart illustrating exemplary operations of a head-worn electronic device for determining user preference information.

[0065] Referring to FIG. 6, according to one embodiment, in operation 600, at least one processor (210) can add data regarding an object into a database. Operation 600 may correspond to operations 330 and 350 of FIG. 3.

[0066] According to one embodiment, in operation 610, at least one processor (210) can identify whether the amount of data in the database is greater than a reference amount. For example, the reference amount may be described as the amount of data in the database required to determine user preference information. For example, the reference amount may be predetermined (or pre-set) or changed (or set) by the user. For example, the amount of data may be described as the number of embedding vectors in the vector database (or vector store).

[0067] According to one embodiment, in operation 620, at least one processor (210) may determine user preference information using a database based on identifying that the amount of data in the database is greater than a reference amount. For example, user preference information may be referred to as perspective information. For example, user preference information may be provided to a model trained with prompts to refine responses regarding prompts. Providing user preference information to a model trained with prompts may be defined as prompting (or prompt design). For example, the trained model may include the generative AI (artificial intelligence) model of FIG. 12. For example, at least one processor (210) may learn the trained model by providing user preference information to the trained model. For example, the learning of the trained model may include supervised learning, unsupervised learning, and / or reinforcement learning. For example, user preference information may be changed (or updated) as data regarding the object is added to the database.

[0068] According to one embodiment, in operation 630, at least one processor (210) may compare a database with another database based on identifying that the amount of data in the database is smaller than a reference amount. For example, because the amount of data in the database is smaller than a reference amount, the amount of data required to determine user preference information may be insufficient. For example, the other database may include a database stored within a head-wearing electronic device (200) and a database learned by a trained model within the head-wearing electronic device (200). For example, the other database may include a database stored within an external electronic device and a database learned by a trained model within the external electronic device.

[0069] According to one embodiment, at least one processor (210) can determine data from another database to be added to the database by comparing the database with another database. For example, data required in the database to determine user preference information can be determined as data from another database to be added to the database.

[0070] According to one embodiment, in operation 640, at least one processor (210) can add data from another database to the database. For example, at least some of the data from another database can be added to the database. For example, at least one processor (210) can increase the amount of data in the database to a reference amount by adding data from another database to the database.

[0071] According to another embodiment, at least one processor (210) can use a trained model to identify missing data to determine user preference information. Based on identifying missing data to determine user preference information, at least one processor (210) can use a trained model to acquire said data. At least one processor (210) can add said data to a database.

[0072] According to one embodiment, in operation 650, at least one processor (210) can determine user preference information by using a database to which data from another database has been added. Refer to the description of operation 620 for determining user preference information.

[0073] According to one embodiment, the database may be a vector database (or vector store), and the data within the database may be embedding vectors. Adding embedding vectors into the vector database is exemplified in the description of FIG. 7.

[0074] Figure 7 illustrates a vector database and other vector databases.

[0075] Referring to FIG. 7, according to one embodiment, the database is a vector database (or vector store) (500), and the data of the database may be embedding vectors (505). Another database is another vector database (or other vector store) (700), and the other data of the other database may be embedding vectors (705). FIG. 7 illustrates the database as a vector database (500), the data of the database as embedding vectors (505), another database as another vector database (700), and the data of the other database as embedding vectors (705) for convenience, but this is exemplary and is not limited thereto. The vector database (500) and embedding vectors (505) may be referenced from the vector database (500) and embedding vectors (505) of FIG. 5.

[0076] According to one embodiment, another vector database (700) may include a vector database stored within a head-worn electronic device (200) and a vector database learned by a trained model within the head-worn electronic device (200). For example, the other vector database (700) may include a vector database stored within an external electronic device and a vector database learned by a trained model within the external electronic device.

[0077] According to one embodiment, at least one processor (210) can identify that the number of embedding vectors (505) of the vector database (500) (or embedding vectors (505) included in the vector database (500)) is less than a reference number. Based on identifying that the number of embedding vectors (505) of the vector database (500) is less than a reference number, at least one processor (210) can add at least one embedding vector of another vector database (700) to the vector database (500). At least one processor (210) can supplement the vector database (500) by adding at least one embedding vector of another vector database (700) to the vector database (500). For example, at least one embedding vector added to the vector database (500) can be determined by comparing the vector database (500) with the other vector database (700). For example, at least one processor (210) can identify missing embedding vectors in the vector database (500) by comparing the vector database (500) with another vector database (700). For example, missing embedding vectors in the vector database (500) can be described as embedding vectors to be added to the blank areas of the vector database (500). For example, at least one processor (210) can supplement the embedding vectors corresponding to the missing embedding vectors in the vector database (500) with embedding vectors in the other vector database (700).

[0078] According to one embodiment, at least one processor (210) can identify that the number of embedding vectors (505) in the vector database (500) is greater than a reference number by adding another vector database (700) to the vector database (500). Based on identifying that the number of embedding vectors (505) in the vector database (500) is greater than a reference number, at least one processor (210) can determine user preference information using the vector database (500). For example, user preference information and perspective information may be provided to a model trained with a prompt to refine the response regarding the prompt. Refining the response regarding the prompt by providing user preference information and perspective information to a model trained with a prompt is illustrated in the description of FIG. 8.

[0079] FIG. 8 is a flowchart illustrating exemplary operations of a head-worn electronic device for displaying information for other objects.

[0080] Referring to FIG. 8, according to one embodiment, in operation 800, at least one processor (210) can detect another object corresponding to the object through one or more cameras (240). The object may be described as an object indicated by data added to a database. The other object may be located within the FOV of one or more cameras (240).

[0081] According to one embodiment, in operation 810, at least one processor (210) may generate a prompt regarding another object based on detecting another object corresponding to the object. For example, a trained model may be used to generate a prompt regarding another object. For example, since the trained model is learned using user perspective information, the prompt generated using the trained model may include text requesting the user to perform necessary processing.

[0082] According to another embodiment, at least one processor (210) can display an image containing another object, which is acquired through one or more cameras (240), through a display assembly (230). While displaying the image containing another object, at least one processor (210) can identify user input requesting a user to perform necessary processing regarding the other object. For example, user input may include voice input received through a microphone (260). At least one processor (210) can generate a prompt regarding the other object based on the user input, but is not limited thereto.

[0083] According to one embodiment, in operation 820, at least one processor (210) may provide a prompt regarding another object to a trained model along with user preference information and perspective information. For example, user preference information and perspective information may be provided to the trained model along with the prompt to refine the response regarding the prompt. For example, user preference information and perspective information may include data regarding an object corresponding to another object, since they are determined based on a database. By providing user preference information and perspective information including data regarding an object corresponding to another object to the trained model along with the prompt, the trained model may be caused to provide a response intended by the user (or desired by the user).

[0084] According to one embodiment, in operation 830, at least one processor (210) may obtain information for another object generated by a trained model using prompts, user preference information, and perspective information. For example, information for another object may include text for another object, an image for another object, a video for another object, and / or audio for another object. However, it is not limited thereto. For example, information for another object may include information intended by the user (or desired by the user) because it is generated by a trained model provided with user preference information and perspective information.

[0085] According to one embodiment, in operation 840, at least one processor (210) can display information for another object through a display assembly (230). For example, information for another object may be displayed together with (or superimposed on) an image containing a representation of another object obtained through one or more cameras (240). At least one processor (210) can output information for another object which is audio through a microphone (260). At least one processor (210) can provide information for another object intended by (or desired by) the user by displaying (or outputting) information for another object.

[0086] According to another embodiment, at least one processor (210) may display a window through a display assembly (230) that inquires whether to display information for another object. For example, at least one processor (210) may receive user input to display information for another object through the window that inquires whether to display information for another object. For example, user input to display information for another object may include touch input and voice input for the window that inquires whether to display information for another object. However, it is not limited thereto. For example, at least one processor (210) may display information for another object through the display assembly (230) based on user input to display information for another object. Displaying information for another object is exemplified in the description of FIG. 9.

[0087] Figure 9 illustrates an example of displaying information for other objects.

[0088] Referring to FIG. 9, according to one embodiment, a state (900) may be described as a state in which an object (905) (e.g., a television located relatively far from a head-worn electronic device (200)) is located within the FOV of one or more cameras (240). For example, the object (905) may correspond to an object (e.g., object (415)) regarding data added (or stored) in a database. For example, objects (910-1) and (910-2) may be further located within the FOV of one or more cameras (240). In the state (900), at least one processor (210) can detect the objects (905, 910-1, 910-2) through one or more cameras (240). For example, at least one processor (210) can determine an object of interest among objects (905, 910-1, 910-2) using user recognition information. For example, since the user recognition information is determined using a database containing data regarding an object (e.g., object (415)) corresponding to object (905), object (905) can be determined as the object of interest among objects (905, 910-1, 910-2). At least one processor (210) can generate a prompt regarding the object (905) determined as the object of interest among objects (905, 910-1, 910-2).

[0089] According to one embodiment, at least one processor (210) may provide a prompt regarding an object (905) to a trained model along with user preference information and perspective information. For example, since the user preference information and perspective information are determined using a database containing data regarding an object (e.g., object (415)) corresponding to the object (905), they may include information for processing (e.g., zooming in on the object (905)) that the user (405) wants (or intends).

[0090] According to one embodiment, at least one processor (210) can obtain information (920) for an object (905) generated by a trained model using prompts regarding the object (905) and user preference information and perspective information. For example, the information (920) for the object (905) may include information desired (or intended) by the user (405) (e.g., an enlarged image of the object (905)) because it is generated by a trained model provided with user preference information and perspective information.

[0091] According to one embodiment, based on acquiring information (920) for an object (905), the electronic device (200) may transition from state (900) to state (915). In state (915), at least one processor (210) may display information (920) for the object (905) through a display assembly (230). For example, information (920) for the object (905) may be displayed together with (or superimposed on) an image containing a representation of the object (905) acquired through one or more cameras (240). By displaying information (920) for the object (905), at least one processor (210) may provide information for the object (905) intended by (or desired by) the user. For example, a user can see information (920) for an object (905) displayed through a display assembly (230) and determine whether the information (920) for the object (905) is information that the user wants (or intends). For example, at least one processor (210) may receive user input indicating that the information (920) for the object (905) is not information that the user wants (or intends). For example, the user input may include an input to stop the display of information (920) for the object (905). The actions of the head-worn electronic device (200) in response to the input to stop the display of information (920) for the object (905) are illustrated in the description of FIG. 10.

[0092] FIG. 10 is a flowchart illustrating exemplary operations of a head-worn electronic device for deleting data regarding an object in a database.

[0093] Referring to FIG. 10, according to one embodiment, in operation 1000, at least one processor (210) can display information for another object through a display assembly (230). Operation 1000 may correspond to operation 840 of FIG. 8.

[0094] According to one embodiment, in operation 1010, at least one processor (210) can identify a user input that stops displaying information for another object. For example, a user input that stops displaying information for another object may be described as a user input indicating that the information for another object is not what the user wants (or intended). For example, as the information for another object is not what the user wants (or intended), the user preference information and perspective information used to generate the information for another object may have errors.

[0095] According to one embodiment, in operation 1020, at least one processor (210) may delete data regarding an object corresponding to another object in the database. For example, at least one processor (210) may redetermine user preference information and perspective information based on the database from which data regarding an object corresponding to another object has been deleted. By redetermining user preference information and perspective information, at least one processor (210) may resolve errors in user preference information and perspective information. For example, at least one processor (210) may provide the redetermined user preference information and perspective information to a trained model along with a prompt. At least one processor (210) may redetermine information for another object generated by the trained model using the prompt and the redetermined user information. For example, the redetermined information for another object may include information for another object that the user wants (or intends). For example, as user preference information is redetermined, user preference information may be changed. As user preference information is repeatedly changed, the direction of user preference information may be changed. As the direction of user preference information changes, the perspective information indicating the direction of user preference information may change.

[0096] FIG. 11 is a block diagram of an electronic device in a network environment according to various embodiments.

[0097] Referring to FIG. 11, in a network environment (1100), an electronic device (1101) may communicate with an electronic device (1102) through a first network (1198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (1104) or a server (1108) through a second network (1199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1101) may communicate with the electronic device (1104) through a server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), memory (1130), input module (1150), sound output module (1155), display module (1160), audio module (1170), sensor module (1176), interface (1177), connection terminal (1178), haptic module (1179), camera module (1180), power management module (1188), battery (1189), communication module (1190), subscriber identification module (1196), or antenna module (1197). In some embodiments, at least one of these components (e.g., connection terminal (1178)) may be omitted from the electronic device (1101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).

[0098] The processor (1120) can, for example, execute software (e.g., program (1140)) to control at least one other component (e.g., hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1120) can store commands or data received from other components (e.g., sensor module (1176) or communication module (1190)) in volatile memory (1132), process the commands or data stored in volatile memory (1132), and store the resulting data in non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1101) includes a main processor (1121) and an auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a specified function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as part thereof.

[0099] The auxiliary processor (1123) may control at least some of the functions or states associated with at least one component of the electronic device (1101) (e.g., display module (1160), sensor module (1176), or communication module (1190)) on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1180) or communication module (1190)). According to one embodiment, the auxiliary processor (1123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (1108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0100] The memory (1130) can store various data used by at least one component of the electronic device (1101) (e.g., processor (1120) or sensor module (1176)). The data may include, for example, software (e.g., program (1140)) and input or output data for related commands. The memory (1130) may include volatile memory (1132) or non-volatile memory (1134).

[0101] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).

[0102] The input module (1150) can receive commands or data to be used for a component of the electronic device (1101) (e.g., processor (1120)) from outside the electronic device (1101) (e.g., user). The input module (1150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0103] The sound output module (1155) can output a sound signal to the outside of the electronic device (1101). The sound output module (1155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0104] The display module (1160) can visually provide information to an external (e.g., user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0105] The audio module (1170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150) or output sound through the sound output module (1155) or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (1101).

[0106] The sensor module (1176) can detect the operating state of the electronic device (1101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0107] The interface (1177) may support one or more specified protocols that can be used for the electronic device (1101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1102)). According to one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0108] The connection terminal (1178) may include a connector through which the electronic device (1101) can be physically connected to an external electronic device (e.g., electronic device (1102)). According to one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0109] The haptic module (1179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (1179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0110] The camera module (1180) can capture still images and video. According to one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0111] The power management module (1188) can manage the power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0112] The battery (1189) can supply power to at least one component of the electronic device (1101). According to one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0113] The communication module (1190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may include one or more communication processors that operate independently of the processor (1120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1104) through a first network (1198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) can identify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1196).

[0114] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (1192) can support various requirements specified in the electronic device (1101), external electronic device (e.g., electronic device (1104)), or network system (e.g., second network (1199)). According to one embodiment, the wireless communication module (1192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 114 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0115] An antenna module (1197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1198) or a second network (1199), may be selected from the plurality of antennas, for example, by a communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1197).

[0116] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0117] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0118] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) through a server (1108) connected to a second network (1199). Each of the external electronic devices (1102, or 1104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations performed on the electronic device (1101) may be performed on one or more of the external electronic devices (1102, 1104, or 1108). For example, if the electronic device (1101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1104) or server (1108) may be included within the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0119] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0120] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0121] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0122] Various embodiments of the present document may be implemented as software (e.g., program (1140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (1136) or external memory (1138)) readable by a machine (e.g., electronic device (1101)). For example, a processor (e.g., processor (1120)) of the machine (e.g., electronic device (1101)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0123] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0124] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0125] FIG. 12 illustrates an example of a generative artificial intelligence system according to one embodiment.

[0126] Referring to FIG. 12, the AI ​​system (1200) may include an input / output interface (1210), an AI framework (1220), a generative AI model (1230), and / or a knowledge repository (1290).

[0127] The input / output interface (1210) may receive input. The input may include user input and / or data acquired or generated by an electronic device (e.g., the head-wearing electronic device (200) or electronic device (1101) described above). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (210) or processor (1120)), such as: illuminance data around the electronic device acquired from a sensor or sensor hub (e.g., auxiliary processor (1123); posture data (or orientation data) of the electronic device; temperature inside the electronic device (e.g., display assembly (230)); or temperature of at least one processor (210); size information of the display area of ​​the display assembly (230); and / or images acquired through an image sensor of the electronic device (e.g., included in a camera module (1180)). The user input may include natural language, touch data obtained through a touch circuit included within the display panel (e.g., used to identify input from a finger and / or stylus), an image displayed (and / or to be displayed) on the display panel, and / or video. By example, without limitation, the user input may be received by an input / output interface (1210) along with context information. The context information may be described as additional information obtained in relation to the user input. The context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the context information may include information about one or more software applications executed within the electronic device at the time the user input is received.For example, the above situation information may include information about the location of the electronic device (or the location of the user of the electronic device) when the user input is received. For example, the user input may be integrated with the situation information. For example, the user input with the situation information integrated as the input may be received by the input / output interface (1210).

[0128] The input / output interface (1210) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI ​​system (1200) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.

[0129] The AI ​​framework (1220) can be used to obtain information (or data) about the input from the input / output interface (1210) and to control one or more components related to the AI ​​system (1200) using the obtained information.

[0130] For example, a prompt design component (1221) within an AI framework (1220) can generate or obtain prompts for a generative AI model (1230) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (1221) may be described as an AI component that uses a learning algorithm and / or a neural network to provide prompts that are enhanced over time. For example, the prompt design component (1221) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (1290)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (1230) (e.g., including an LLM or LMM).

[0131] For example, an API / plugin management component (1222) within the AI ​​framework (1220) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (1230). For example, the API / plugin management component (1222) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (1290)). For example, the API / plugin management component (1222) may support access to at least some of the data sources. For example, the API / plugin management component (1222) may be used to request another component (e.g., application / service component (1280)) that performs feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1222) may be provided to the prompt design component (1221) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1222) may be provided to the generative AI model (1230).

[0132] For example, an improvement component (1223) within the AI ​​framework (1220) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1230). For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) is related to the input. For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) contains biased content. For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) contains harmful content. For example, the improvement component (1223) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1230). For example, the improvement component (1223) may support providing a hint to the user to improve the content.

[0133] A generative AI model (1230) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information relative to the prompt, but relative to the prompt. For example, the feedback may include new content relative to the prompt. For example, the generative AI model (1230) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, the generative AI model (1230) may include an LMM that generates the feedback by recognizing characters, images, and / or voice.

[0134] As an example without limitation, the AI ​​framework (1220) and / or generative AI model (1230) may be included within an AI module (e.g., including a processing circuit) within the head-wearing electronic device (200). For example, the AI ​​module may be operatively coupled with at least one processor of the head-wearing electronic device (200) (e.g., at least one processor (210) or processor (1120)). For example, the AI ​​module may be operatively coupled with a display driving circuit of the electronic device. For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0135] FIG. 13a shows an example of a perspective view of a wearable device.

[0136] FIG. 13b shows an example of one or more hardware components placed within a wearable device.

[0137] FIG. 13a illustrates an example of a perspective view of a wearable device. FIG. 13b illustrates an example of one or more hardware components disposed within the wearable device. According to one embodiment, a head-wearing electronic device (200) may have the form of glasses that are wearable on a part of a user's body (e.g., head). The head-wearing electronic device (200) of FIG. 13a and FIG. 13b may be an example of the head-wearing electronic device (200) of FIG. 2. The head-wearing electronic device (200) may include a head-mounted display (HMD). For example, the housing of the head-wearing electronic device (200) may include a flexible material such as rubber and / or silicone that has a shape that adheres to a part of the user's head (e.g., a part of the face covering both eyes). For example, the housing of the head-worn electronic device (200) may include one or more straps that can be twined around the user's head, and / or one or more temples that can be attached to the ears of the head.

[0138] Referring to FIG. 13a, a head-wearing electronic device (200) according to one embodiment may include at least one display (1350) and a frame (1300) supporting at least one display (1350).

[0139] According to one embodiment, a head-worn electronic device (200) may be worn on a part of a user's body. The head-worn electronic device (200) may provide augmented reality (AR), virtual reality (VR), or mixed reality (MR) that combines augmented reality and virtual reality to a user wearing the head-worn electronic device (200). For example, the head-worn electronic device (200) may display a virtual reality image provided by at least one optical device (1382, 1384) of FIG. 13b on at least one display (1350) in response to a designated gesture of the user obtained through the motion recognition camera (1360-2, 1360-3) of FIG. 13b.

[0140] According to one embodiment, at least one display (1350) can provide visual information to a user. For example, at least one display (1350) may include a transparent or translucent lens. At least one display (1350) may include a first display (1350-1) and / or a second display (1350-2) spaced apart from the first display (1350-1). For example, the first display (1350-1) and the second display (1350-2) may be positioned at locations corresponding to the user's left eye and right eye, respectively.

[0141] Referring to FIG. 13b, at least one display (1350) may provide visual information transmitted from external light to a user through a lens included in at least one display (1350) and other visual information distinct from said visual information. The lens may be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens. For example, at least one display (1350) may include a first surface (1331) and a second surface (1332) opposite to the first surface (1331). A display area may be formed on the second surface (1332) of at least one display (1350). When a user wears the head-worn electronic device (200), external light may be transmitted to the user by being incident on the first surface (1331) and transmitted through the second surface (1332). As another example, at least one display (1350) can display an augmented reality image combined with a virtual reality image provided by at least one optical device (1382, 1384) on a real image transmitted through external light in a display area formed on a second surface (1332).

[0142] In one embodiment, at least one display (1350) may include at least one waveguide (1333, 1334) that diffracts light emitted from at least one optical device (1382, 1384) and transmits it to a user. At least one waveguide (1333, 1334) may be formed based on at least one of glass, plastic, or polymer. A nano pattern may be formed on the exterior or at least a portion of the interior of at least one waveguide (1333, 1334). The nano pattern may be formed based on a polygonal and / or curved grating structure. Light incident on one end of at least one waveguide (1333, 1334) may be propagated to the other end of at least one waveguide (1333, 1334) by the nano pattern. At least one waveguide (1333, 1334) may include at least one diffractive element (e.g., DOE (diffractive optical element), HOE (holographic optical element)) and at least one reflective element (e.g., a reflective mirror). For example, at least one waveguide (1333, 1334) may be placed within a head-worn electronic device (200) to guide a screen displayed by at least one display (1350) to the user's eye. For example, the screen may be transmitted to the user's eye based on total internal reflection (TIR) ​​occurring within at least one waveguide (1333, 1334).

[0143] The head-wearing electronic device (200) can analyze objects included in real-world images collected through a camera (1360-4), combine virtual objects corresponding to objects among the analyzed objects that are the target of augmented reality provision, and display them on at least one display (1350). The virtual object may include at least one of text and images regarding various information related to the objects included in the real-world images. The head-wearing electronic device (200) can analyze objects based on multi-cameras, such as stereo cameras. For the object analysis, the head-wearing electronic device (200) can perform spatial recognition (e.g., SLAM (simultaneous localization and mapping)) using multi-cameras and / or time-of-flight (ToF). A user wearing the head-wearing electronic device (200) can view images displayed on at least one display (1350).

[0144] According to one embodiment, the frame (1300) may be formed as a physical structure that allows the head-worn electronic device (200) to be worn on the user's body. According to one embodiment, the frame (1300) may be configured such that when the user wears the head-worn electronic device (200), the first display (1350-1) and the second display (1350-2) can be positioned corresponding to the user's left and right eyes. The frame (1300) may support at least one display (1350). For example, the frame (1300) may support the first display (1350-1) and the second display (1350-2) so that they are positioned corresponding to the user's left and right eyes.

[0145] Referring to FIG. 13a, the frame (1300) may include an area (1320) in which at least a portion of the frame contacts a portion of the user's body when the user wears the head-worn electronic device (200). For example, the area (1320) of the frame (1300) in contact with a portion of the user's body may include an area in contact with a portion of the user's nose, a portion of the user's ear, and a portion of the side of the user's face that the head-worn electronic device (200) contacts. According to one embodiment, the frame (1300) may include a nose pad (1310) that contacts a portion of the user's body. When the head-worn electronic device (200) is worn by the user, the nose pad (1310) may contact a portion of the user's nose. The frame (1300) may include a first temple (1304) and a second temple (1305) that come into contact with another part of the user's body that is distinct from the part of the user's body.

[0146] For example, the frame (1300) may include a first rim (1301) covering at least a portion of a first display (1350-1), a second rim (1302) covering at least a portion of a second display (1350-2), a bridge (1303) positioned between the first rim (1301) and the second rim (1302), a first pad (1311) positioned along a portion of the edge of the first rim (1301) from one end of the bridge (1303), a second pad (1312) positioned along a portion of the edge of the second rim (1302) from the other end of the bridge (1303), a first temple (1304) extending from the first rim (1301) and fixed to a portion of the wearer's ear, and a second temple (1305) extending from the second rim (1302) and fixed to a portion of the ear opposite to the ear. The first pad (1311) and the second pad (1312) may come into contact with a part of the user's nose, and the first temple (1304) and the second temple (1305) may come into contact with a part of the user's face and a part of the ear. The temples (1304, 1305) may be rotatably connected to the rim through the hinge units (1306, 1307) of FIG. 13b. The first temple (1304) may be rotatably connected to the first rim (1301) through a first hinge unit (1306) positioned between the first rim (1301) and the first temple (1304). The second temple (1305) may be rotatably connected to the second rim (1302) through a second hinge unit (1307) disposed between the second rim (1302) and the second temple (1305). According to one embodiment, the head-wearing electronic device (200) may identify an external object (e.g., a user's fingertip) touching the frame (1300) and / or a gesture performed by said external object by using a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of the surface of the frame (1300).

[0147] According to one embodiment, the head-wearing electronic device (200) may include hardware that performs various functions (e.g., hardware described below based on the block diagram of FIG. 13). For example, the hardware may include a battery module (1370), an antenna module (1375), at least one optical device (1382, 1384), speakers (e.g., speakers (1355-1, 1355-2)), microphones (e.g., microphones (1365-1, 1365-2, 1365-3)), a light-emitting module (not shown), and / or a printed circuit board (1390) (e.g., a printed circuit board). The various hardware may be placed within a frame (1300).

[0148] According to one embodiment, a microphone (e.g., microphones (1365-1, 1365-2, 1365-3)) of a head-wearing electronic device (200) is positioned on at least a portion of a frame (1300) to acquire a sound signal. A first microphone (1365-1) positioned on a bridge (1303), a second microphone (1365-2) positioned on a second rim (1302), and a third microphone (1365-3) positioned on a first rim (1301) are shown in FIG. 13b, but the number and position of the microphones (1365) are not limited to the embodiment of FIG. 13b. If the number of microphones (1365) included in the head-wearing electronic device (200) is two or more, the head-wearing electronic device (200) can identify the direction of a sound signal by using multiple microphones placed on different parts of the frame (1300).

[0149] According to one embodiment, at least one optical device (1382, 1384) may project a virtual object onto at least one display (1350) to provide various image information to a user. For example, at least one optical device (1382, 1384) may be a projector. At least one optical device (1382, 1384) may be disposed adjacent to at least one display (1350) or included within at least one display (1350) as part of at least one display (1350). According to one embodiment, a head-worn electronic device (200) may include a first optical device (1382) corresponding to a first display (1350-1) and a second optical device (1384) corresponding to a second display (1350-2). For example, at least one optical device (1382, 1384) may include a first optical device (1382) positioned at the edge of a first display (1350-1) and a second optical device (1384) positioned at the edge of a second display (1350-2). The first optical device (1382) may transmit light to a first waveguide (1333) positioned on the first display (1350-1), and the second optical device (1384) may transmit light to a second waveguide (1334) positioned on the second display (1350-2).

[0150] In one embodiment, the camera (1360) may include a shooting camera (1360-4), an eye tracking camera (ET CAM) (1360-1), and / or a motion recognition camera (1360-2, 1360-3). The shooting camera (1360-4), the eye tracking camera (1360-1), and the motion recognition camera (1360-2, 1360-3) may be positioned at different locations on the frame (1300) and may perform different functions. The eye tracking camera (1360-1) may output data indicating the position of the eyes or the gaze of a user wearing the head-worn electronic device (200). For example, the head-worn electronic device (200) may detect the gaze from an image containing the user's pupils obtained through the eye tracking camera (1360-1). The head-wearing electronic device (200) can identify an object focused by the user (e.g., a real object, and / or a virtual object) by using the user's gaze acquired through the eye-tracking camera (1360-1). The head-wearing electronic device (200), having identified the focused object, can execute a function (e.g., gaze interaction) for interaction between the user and the focused object. The head-wearing electronic device (200) can represent a portion corresponding to the eyes of an avatar representing the user in a virtual space by using the user's gaze acquired through the eye-tracking camera (1360-1). The head-wearing electronic device (200) can render an image (or screen) displayed on at least one display (1350) based on the position of the user's eyes. For example, the visual quality of a first region associated with the gaze within the image and the visual quality of a second region distinct from the first region (e.g., resolution, brightness, saturation, grayscale, PPI) may differ from each other.The head-worn electronic device (200) can acquire an image having a visual quality of a first region and a visual quality of a second region that matches the user's gaze by using foveated rendering. For example, if the head-worn electronic device (200) supports an iris recognition function, user authentication can be performed based on iris information acquired using an eye-tracking camera (1360-1). An example in which the eye-tracking camera (1360-1) is positioned toward the user's right eye is shown in FIG. 13b, but the embodiment is not limited thereto, and the eye-tracking camera (1360-1) may be positioned alone toward the user's left eye or toward both eyes.

[0151] In one embodiment, the camera (1360-4) can capture a real image or background to be matched with a virtual image in order to implement augmented reality or mixed reality content. The camera (1360-4) can be used to acquire high-resolution images based on HR (high resolution) or PV (photo video). The camera (1360-4) can capture an image of a specific object located at the position viewed by the user and provide the image to at least one display (1350). The at least one display (1350) can display a single image in which information regarding a real image or background including the image of the specific object acquired using the camera (1360-4) and a virtual image provided through at least one optical device (1382, 1384) are superimposed. The head-wearing electronic device (200) can compensate for depth information (e.g., the distance between the head-wearing electronic device (200) and an external object obtained through a depth sensor) using an image obtained through a shooting camera (1360-4). The head-wearing electronic device (200) can perform object recognition using an image obtained through a shooting camera (1360-4). The head-wearing electronic device (200) can perform a function of focusing on an object (or subject) within an image (e.g., auto focus) and / or an optical image stabilization (OIS) function (e.g., anti-shake function) using a shooting camera (1360-4). The head-wearing electronic device (200) can perform a pass-through function to display an image obtained through a shooting camera (1360-4) superimposed on at least a portion of a screen representing a virtual space while displaying a screen representing a virtual space on at least one display (1350).In one embodiment, the camera (1360-4) may be placed on a bridge (1303) positioned between the first rim (1301) and the second rim (1302).

[0152] The eye tracking camera (1360-1) can achieve more realistic augmented reality by tracking the gaze of a user wearing a head-worn electronic device (200), thereby matching the user's gaze with visual information provided to at least one display (1350). For example, the head-worn electronic device (200) can naturally display environmental information related to the user's front at the location where the user is situated on at least one display (1350) when the user looks straight ahead. The eye tracking camera (1360-1) may be configured to capture an image of the user's pupil to determine the user's gaze. For example, the eye tracking camera (1360-1) may receive a gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In one embodiment, the eye tracking camera (1360-1) may be positioned at locations corresponding to the user's left and right eyes. For example, the eye-tracking camera (1360-1) may be positioned within the first rim (1301) and / or the second rim (1302) to face the direction in which the user wearing the head-worn electronic device (200) is located.

[0153] A motion recognition camera (1360-2, 1360-3) can provide a specific event to a screen provided on at least one display (1350) by recognizing the movement of the user's entire body or part thereof, such as the user's torso, hands, or face. A motion recognition camera (1360-2, 1360-3) can recognize the user's gesture, acquire a signal corresponding to the gesture, and provide a display corresponding to the signal to at least one display (1350). A processor can identify the signal corresponding to the gesture and, based on the identification, perform a designated function. A motion recognition camera (1360-2, 1360-3) can be used to perform spatial recognition functions using SLAM and / or depth maps for a 6-degrees-of-freedom pose (6 dof pose). A processor can use the motion recognition camera (1360-2, 1360-3) to perform gesture recognition functions and / or object tracking functions. In one embodiment, a motion recognition camera (1360-2, 1360-3) may be placed on the first rim (1301) and / or the second rim (1302).

[0154] The camera (1360) included in the head-worn electronic device (200) is not limited to the eye-tracking camera (1360-1) and motion recognition camera (1360-2, 1360-3) described above. For example, the head-worn electronic device (200) can identify external objects included within the FoV by using a camera positioned toward the user's FoV. The identification of external objects by the head-worn electronic device (200) can be performed based on a sensor for identifying the distance between the head-worn electronic device (200) and the external object, such as a depth sensor and / or a time of flight (ToF) sensor. The camera (1360) positioned toward the FoV may support an autofocus function and / or an optical image stabilization (OIS) function. For example, the head-worn electronic device (200) may include a camera (1360) (e.g., a face tracking camera) positioned toward the face to acquire an image including the face of a user wearing the head-worn electronic device (200).

[0155] Although not illustrated, according to one embodiment, a head-worn electronic device (200) may further include a light source (e.g., LED) that emits light toward a subject (e.g., user's eyes, face, and / or external objects within the FoV) being photographed using a camera (1360). The light source may include an LED of infrared wavelength. The light source may be placed in at least one of a frame (1300) and hinge units (1306, 1307).

[0156] According to one embodiment, the battery module (1370) can supply power to the electronic components of the head-wearing electronic device (200). In one embodiment, the battery module (1370) may be placed within the first temple (1304) and / or the second temple (1305). For example, the battery module (1370) may be a plurality of battery modules (1370). The plurality of battery modules (1370) may each be placed in the first temple (1304) and the second temple (1305), respectively. In one embodiment, the battery module (1370) may be placed at the end of the first temple (1304) and / or the second temple (1305).

[0157] The antenna module (1375) can transmit a signal or power to the outside of the head-wearing electronic device (200) or receive a signal or power from the outside. In one embodiment, the antenna module (1375) may be placed within the first temple (1304) and / or the second temple (1305). For example, the antenna module (1375) may be placed near one side of the first temple (1304) and / or the second temple (1305).

[0158] A speaker (1355) can output an acoustic signal to the outside of the head-wearing electronic device (200). The acoustic output module may be referred to as a speaker. In one embodiment, the speaker (1355) may be placed within a first temple (1304) and / or a second temple (1305) to be positioned adjacent to the ears of a user wearing the head-wearing electronic device (200). For example, the speaker (1355) may include a second speaker (1355-2) positioned adjacent to the user's left ear by being placed within the first temple (1304), and a first speaker (1355-1) positioned adjacent to the user's right ear by being placed within the second temple (1305).

[0159] A light-emitting module (not shown) may include at least one light-emitting element. The light-emitting module may emit light of a color corresponding to a specific state or emit light with an action corresponding to a specific state in order to visually provide information regarding a specific state of the head-wearing electronic device (200) to the user. For example, if the head-wearing electronic device (200) requires charging, it may emit red light at a constant frequency. In one embodiment, the light-emitting module may be placed on the first rim (1301) and / or the second rim (1302).

[0160] Referring to FIG. 13b, a head-wearing electronic device (200) according to one embodiment may include a printed circuit board (PCB) (1390). The PCB (1390) may be included in at least one of a first temple (1304) or a second temple (1305). The PCB (1390) may include an interposer disposed between at least two sub-PCBs. On the PCB (1390), one or more hardware components included in the head-wearing electronic device (200) (e.g., hardware components illustrated by different blocks in FIG. 4) may be disposed. The head-wearing electronic device (200) may include a flexible PCB (FPCB) for interconnecting the hardware components.

[0161] According to one embodiment, a head-worn electronic device (200) may include at least one of a gyroscope sensor, a gravity sensor, and / or an acceleration sensor for detecting the posture of the head-worn electronic device (200) and / or the posture of a body part (e.g., head) of a user wearing the head-worn electronic device (200). Each of the gravity sensor and the acceleration sensor may measure gravitational acceleration and / or acceleration based on designated three-dimensional axes (e.g., x-axis, y-axis, and z-axis) that are perpendicular to each other. The gyroscope sensor may measure the angular velocity of each of the designated three-dimensional axes (e.g., x-axis, y-axis, and z-axis). At least one of the gravity sensor, the acceleration sensor, and the gyroscope sensor may be referred to as an inertial measurement unit (IMU). According to one embodiment, the head-wearing electronic device (200) can identify a user's motion and / or gesture performed to execute or interrupt a specific function of the head-wearing electronic device (200) based on an IMU.

[0162] FIGS. 14a and 14b show examples of the appearance of a wearable device.

[0163] FIGS. 14a and 14b illustrate an example of the appearance of a wearable device (e.g., a head-wearing electronic device (200)). The head-wearing electronic device (200) of FIGS. 14a and 14b may be an example of the head-wearing electronic device (200) of FIG. 2. According to one embodiment, an example of the appearance of a first surface (1410) of the housing of the head-wearing electronic device (200) may be illustrated in FIG. 14a, and an example of the appearance of a second surface (1420) opposite to the first surface (1410) may be illustrated in FIG. 14b.

[0164] Referring to FIG. 14a, according to one embodiment, a first surface (1410) of a head-wearing electronic device (200) may have a shape that is attachable to a part of a user's body (e.g., the user's face). Although not illustrated, the head-wearing electronic device (200) may further include a strap for securing to a part of a user's body and / or one or more temples (e.g., a first temple (1304) and / or a second temple (1305) of FIG. 13a to FIG. 13b). A first display (1350-1) for outputting an image to the left eye among the user's two eyes, and a second display (1350-2) for outputting an image to the right eye among the two eyes may be disposed on the first surface (1410). The head-wearing electronic device (200) may further include rubber or silicone packing formed on the first surface (1410) to prevent interference by light different from light emitted from the first display (1350-1) and the second display (1350-2) (e.g., ambient light).

[0165] According to one embodiment, a head-worn electronic device (200) may include cameras (1360-1) for photographing and / or tracking both eyes of a user adjacent to each of the first display (1350-1) and the second display (1350-2). The cameras (1360-1) may be referenced to the eye-tracking camera (1360-1) of FIG. 13b. According to one embodiment, a head-worn electronic device (200) may include cameras (1360-5, 1360-6) for photographing and / or recognizing a user's face. The cameras (1360-5, 1360-6) may be referenced to FT cameras. The head-worn electronic device (200) may control an avatar representing the user in a virtual space based on the motion of the user's face identified using the cameras (1360-5, 1360-6). For example, the head-worn electronic device (200) can change the texture and / or shape of a part of an avatar (e.g., a part of an avatar representing a human face) by using information obtained by cameras (1360-5, 1360-6) (e.g., FT cameras) and representing the facial expression of a user wearing the head-worn electronic device (200).

[0166] Referring to FIG. 14b, on a second surface (1420) opposite to the first surface (1410) of FIG. 14a, a camera (e.g., cameras (1360-7, 1360-8, 1360-9, 1360-10, 1360-11, 1360-12)), and / or a sensor (e.g., a depth sensor (1430)) may be placed to obtain information related to the external environment of the head-wearing electronic device (200). For example, cameras (1360-7, 1360-8, 1360-9, 1360-10) may be placed on the second surface (1420) to recognize external objects. The cameras (1360-7, 1360-8, 1360-9, 1360-10) can be referenced to the motion recognition cameras (1360-2, 1360-3) of FIG. 13b.

[0167] For example, using cameras (1360-11, 1360-12), the head-wearing electronic device (200) can acquire images and / or videos to be transmitted to each of the user's two eyes. Camera (1360-11) may be placed on the second surface (1420) of the head-wearing electronic device (200) to acquire an image to be displayed through a second display (1350-2) corresponding to the right eye among the two eyes. Camera (1360-12) may be placed on the second surface (1420) of the head-wearing electronic device (200) to acquire an image to be displayed through a first display (1350-1) corresponding to the left eye among the two eyes. Cameras (1360-11, 1360-12) may be referenced to the shooting camera (1360-4) of FIG. 13b.

[0168] According to one embodiment, a head-worn electronic device (200) may include a depth sensor (1430) disposed on a second surface (1420) to identify the distance between the head-worn electronic device (200) and an external object. Using the depth sensor (1430), the head-worn electronic device (200) may acquire spatial information (e.g., a depth map) for at least a portion of the FoV of a user wearing the head-worn electronic device (200). Although not illustrated, a microphone may be disposed on the second surface (1420) of the head-worn electronic device (200) to acquire sound output from an external object. The number of microphones may be one or more, depending on the embodiment.

[0169] Hereinafter, with reference to FIG. 15, the hardware or software configuration of the head-worn electronic device (200) is described.

[0170] Figure 15 shows an example of a block diagram of a wearable device.

[0171] FIG. 15 illustrates an example of a block diagram of a wearable device (e.g., a head-worn electronic device (200)). The head-worn electronic device (200) of FIG. 15 may be an example of the head-worn electronic device (200) of FIG. 2 and the head-worn electronic device (200) of FIG. 13a through FIG. 14b.

[0172] Referring to FIG. 15, a head-wearing electronic device (200) according to one embodiment may include a processor (1510), memory (1515), a display (1350) (e.g., a first display (1350-1) and / or a second display (1350-2) of FIG. 13a, FIG. 13b, FIG. 14a, and FIG. 14b), and / or a sensor. The processor (1510), memory (1515), display (1350), and / or sensor may be electrically and / or operationally connected to each other by an electronic component such as a communication bus (1502). In the present disclosure, the operational connection of the electronic components may include a direct connection established between the electronic components and / or an indirect connection established between the electronic components so that a first electronic component among the electronic components is controlled by a second electronic component among the electronic components. The type and / or number of electronic components included in the head-wearing electronic device (200) are not limited to those shown in FIG. 15. For example, the head-wearing electronic device (200) may include only some of the electronic components shown in FIG. 15.

[0173] A processor (1510) of a head-wearing electronic device (200) according to one embodiment may include a circuit (e.g., a processing circuit) for processing data based on one or more instructions. The circuit for processing data may include, for example, an arithmetic and logic unit (ALU), a field programmable gate array (FPGA), a central processing unit (CPU), and / or an application processor (AP). In one embodiment, the head-wearing electronic device (200) may include one or more processors. The processor (1510) may have a structure of a multi-core processor such as a dual core, a quad core, a hexa core, and / or an octa core. The multi-core processor structure of the processor (1510) may include a structure based on multiple core circuits (e.g., a big-little structure), distinguished by power consumption, clock, and / or computational power per unit time. In one embodiment comprising a processor (1510) having a multi-core processor structure, the operations and / or functions of the present disclosure may be performed individually or collectively by one or more cores included in the processor (1510).

[0174] A memory (1515) of a head-worn electronic device (200) according to one embodiment may include electronic components for storing data and / or instructions that are input to or output from a processor (1510). The memory (1515) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, and embedded multimedia card (eMMC). In one embodiment, the memory (1515) may be referred to as storage.

[0175] In one embodiment, a display (1350) of a head-wearing electronic device (200) can output visualized information to a user of the head-wearing electronic device (200). A display (1350) arranged in front of the eyes of a user wearing the head-wearing electronic device (200) may be placed in at least a part of the housing of the head-wearing electronic device (200) (e.g., a first display (1350-1) and / or a second display (1350-2) of FIG. 13a, FIG. 13b, FIG. 14a, and FIG. 14b). For example, the display (1350) may be controlled by a processor (1510) including circuits such as a CPU, a GPU (graphic processing unit), and / or a DPU (display processing unit) to output visualized information to the user. The display (1350) may include a flexible display, a flat panel display (FPD), and / or electronic paper. The display (1350) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). Embodiments are not limited thereto, for example, if the head-wearing electronic device (200) includes a lens for transmitting external light (or ambient light), the display (1350) may include a projector (or projection assembly) for projecting light onto said lens. In one embodiment, the display (1350) may be referred to as a display panel and / or a display module. Pixels included in the display (1350) may be positioned toward either of the user's two eyes when worn by the user of the head-wearing electronic device (200).For example, the display (1350) may include display areas (or active areas) corresponding to each of the user's two eyes.

[0176] In one embodiment, a sensor of the head-wearing electronic device (200) may generate electrical information that can be processed by a processor (1510) and / or a memory (1515) from non-electronic information associated with the head-wearing electronic device (200). For example, the sensor may include a global positioning system (GPS) sensor for detecting the geographic location of the head-wearing electronic device (200). In addition to the GPS method, the sensor may generate information indicating the geographic location of the head-wearing electronic device (200) based on a global navigation satellite system (GNSS), such as Galileo or Beidou (compass). The information may be stored in the memory (1515), processed by the processor (1510), and / or transmitted to another electronic device distinct from the head-wearing electronic device (200) via a communication circuit.

[0177] According to one embodiment, within the memory (1515) of the head-wearing electronic device (200), one or more instructions (or commands) representing data to be processed by the processor (1510) of the head-wearing electronic device (200), calculations to be performed, and / or operations may be stored. A set of one or more instructions may be referred to as a program, firmware, operating system, process, routine, sub-routine, and / or software application (hereinafter, application). For example, the head-wearing electronic device (200), and / or processor (1510) may perform at least one of the operations of FIGS. 3 through 10 when a set of a plurality of instructions distributed in the form of an operating system, firmware, driver, program, and / or software application is executed. In the following, the statement that a software application is installed within a head-worn electronic device (200) may mean that one or more instructions provided in the form of a software application (or package) are stored in memory (1515), and that said one or more applications are stored in an executable format (e.g., a file having an extension specified by the operating system of the head-worn electronic device (200)) by the processor (1510). For example, the application may include a program and / or library related to a service provided to a user.

[0178] Referring to FIG. 15, programs installed on a head-worn electronic device (200) may be included in any one of different layers, including an application layer (1540), a framework layer (1550), and / or a hardware abstraction layer (HAL) (1580), depending on the target. For example, within the hardware abstraction layer (1580), programs (e.g., modules, or drivers) designed to target the hardware of the head-worn electronic device (200) (e.g., a display (1350), and / or sensors) may be included. The framework layer (1550) may be referred to as an XR framework layer in that it includes one or more programs for providing XR (extended reality) services. For example, the layers illustrated in FIG. 15 are logically (or for convenience of explanation) separated, and this does not mean that the address space of memory (1515) is separated by said layers.

[0179] For example, within the framework layer (1550), programs designed to target at least one of the hardware abstraction layer (1580) and / or the application layer (1540) (e.g., a location tracker (1571), a spatial recognizer (1572), a gesture tracker (1573), an eye tracker (1574), and / or a face tracker (1575)) may be included. The programs included in the framework layer (1550) may provide an application programming interface (API) that is executable (or invokeable) based on other programs.

[0180] For example, within the application layer (1540), a program designed to target a user of the head-worn electronic device (200) may be included. Examples of programs included in the application layer (1540) include an XR (extended reality) system UI (user interface) (1541) and / or an XR application (1542), but embodiments are not limited thereto. For example, programs included in the application layer (1540) (e.g., software applications) may call an API to cause the execution of a function supported by programs included in the framework layer (1550).

[0181] For example, the head-wearing electronic device (200) may display one or more visual objects on the display (1350) to perform interaction with the user based on the execution of the XR system UI (1541). A visual object may mean an object that can be placed on the screen for the transmission of information and / or interaction, such as text, images, icons, videos, buttons, checkboxes, radio buttons, text boxes, sliders, and / or tables. A visual object may be referred to as a visual guide, a virtual object, a visual element, a UI element, a view object, and / or a view element. The head-wearing electronic device (200) may provide the user with functions available in a virtual space based on the execution of the XR system UI (1541).

[0182] Referring to FIG. 15, a lightweight renderer (1543) and / or an XR plugin (1544) are depicted within the XR system UI (1541), but are not limited thereto. For example, based on the XR system UI (1541), the processor (1510) may execute a lightweight renderer (1543) and / or an XR plugin (1544) within the framework layer (1550).

[0183] For example, a head-worn electronic device (200) may acquire resources (e.g., APIs, system processes and / or libraries) used to define, create, and / or execute a rendering pipeline, which is permitted to be partially modified, based on the execution of a lightweight renderer (1543). The lightweight renderer (1543) may be referred to as a lightweight render pipeline in terms of defining a rendering pipeline, which is permitted to be partially modified. The lightweight renderer (1543) may include a renderer built prior to the execution of a software application (e.g., a pre-built renderer). For example, the head-worn electronic device (200) may acquire resources (e.g., APIs, system processes and / or libraries) used to define, create, and / or execute the entire rendering pipeline based on the execution of an XR plugin (1544). The XR plugin (1544) can be referred to as an open XR native client in terms of defining (or setting) the entire rendering pipeline.

[0184] For example, the head-worn electronic device (200) may display a screen representing at least a portion of a virtual space on a display (1350) based on the execution of an XR application (1542). An XR plugin (1544-1) included in the XR application (1542) may include instructions that support functions similar to those of an XR plugin (1544) of an XR system UI (1541). Descriptions of the XR plugin (1544-1) that overlap with descriptions of the XR plugin (1544) may be omitted. The head-worn electronic device (200) may trigger the execution of a virtual space manager (1551) based on the execution of the XR application (1542).

[0185] For example, the head-wearing electronic device (200) may display an image on a display (1350) in virtual space based on the execution of an application (1545). The application (1545) may be configured to output image information for displaying a two-dimensional image. The head-wearing electronic device (200) may trigger the execution of a virtual space manager (1551) based on the execution of the application (1545). The head-wearing electronic device (200) may generate dual image information to display the two-dimensional image in three-dimensional virtual space based on the execution of the application (1545). Here, the dual image information may include a first image information for the left eye and a second image information for the right eye, taking into account binocular parallax. To display the two-dimensional image in three-dimensional virtual space, the head-wearing electronic device (200) may generate the dual image information based on the image information for displaying the two-dimensional image.

[0186] According to one embodiment, the head-worn electronic device (200) may provide virtual space services based on the execution of a virtual space manager (1551). For example, the virtual space manager (1551) may include a platform for supporting virtual space services. Based on the execution of the virtual space manager (1551), the head-worn electronic device (200) may identify a virtual space formed based on the user's location indicated by data acquired through a sensor (1430) and may display at least a portion of the virtual space on a display (1350). The virtual space manager (1551) may be referred to as a composition presentation manager (CPM).

[0187] For example, the virtual space manager (1551) may include a runtime service (1552). For example, the runtime service (1552) may be referred to as an OpenXR runtime module (or OpenXR runtime program). The head-worn electronic device (200) may execute at least one of a user pose prediction function, a frame timing function, and / or a spatial input function based on the execution of the runtime service (1552). For example, the head-worn electronic device (200) may perform rendering for a virtual space service for the user based on the execution of the runtime service (1552). For example, a virtual space-related function executable by the application layer (1540) may be supported based on the execution of the runtime service (1552).

[0188] For example, the virtual space manager (1551) may include a pass-through manager (1553). The head-worn electronic device (200) may display an image and / or video representing the real space acquired through an external camera superimposed on at least a portion of the screen while displaying a screen representing the virtual space on the display (1350) based on the execution of the pass-through manager (1553).

[0189] For example, the virtual space manager (1551) may include an input manager (1554). The head-worn electronic device (200) may identify acquired data (e.g., sensor data) by executing one or more programs included within the recognition service layer (1570) based on the execution of the input manager (1554). The head-worn electronic device (200) may identify user inputs associated with the head-worn electronic device (200) using the acquired data. The user inputs may be associated with user motions (e.g., hand gestures), gaze, and / or speech identified by a sensor (e.g., an image sensor (1430) such as an external camera). The user inputs may be identified based on an external electronic device connected (or paired) via a communication circuit.

[0190] For example, the perception abstract layer (1560) can be used for data exchange between the virtual space manager (1551) and the perception service layer (1570). In terms of being used for data exchange between the virtual space manager (1551) and the perception service layer (1570), the perception abstract layer (1560) can be referred to as an interface. As an example, the perception abstract layer (1560) can be referred to as OpenPX. The perception abstract layer (1560) can be used for a perception client and a perception service.

[0191] According to one embodiment, the recognition service layer (1570) may include one or more programs for processing data acquired from a sensor. The one or more programs may include at least one of a location tracker (1571), a spatial recognizer (1572), a gesture tracker (1573), and / or an eye tracker (1574). The type and / or number of the one or more programs included in the recognition service layer (1570) are not limited to those shown in FIG. 15.

[0192] For example, the head-wearing electronic device (200) can identify the posture of the head-wearing electronic device (200) using a sensor (1430) based on the execution of a position tracker (1571). The head-wearing electronic device (200) can identify the 6 degrees of freedom pose (6 dof pose) of the head-wearing electronic device (200) using data acquired using an external camera (e.g., image sensor (1521)) and / or an IMU (e.g., motion sensor (1522) including a gyroscope, accelerometer, and / or geomagnetic sensor) based on the execution of the position tracker (1571). The position tracker (1571) may be referred to as a head tracking (HeT) module (or head tracker, head tracking program).

[0193] For example, the head-wearing electronic device (200) may acquire information to provide a three-dimensional virtual space corresponding to the surrounding environment (e.g., external space) of the head-wearing electronic device (200) (or the user of the head-wearing electronic device (200)) based on the execution of the spatial recognition device (1572). The head-wearing electronic device (200) may reproduce the surrounding environment of the head-wearing electronic device (200) in three dimensions using data acquired using an external camera (e.g., image sensor (1521)) based on the execution of the spatial recognition device (1572). The head-wearing electronic device (200) may identify at least one of a plane, a slope, and a staircase based on the surrounding environment of the head-wearing electronic device (200) reproduced in three dimensions based on the execution of the spatial recognition device (1572). The space recognizer (1572) can be referred to as a scene understanding (SU) module (or scene understanding program).

[0194] For example, the head-wearing electronic device (200) can identify (or recognize) the pose and / or gesture of the user's hand of the head-wearing electronic device (200) based on the execution of the gesture tracker (1573). For example, the head-wearing electronic device (200) can identify the pose and / or gesture of the user's hand using data acquired from an external camera (e.g., image sensor (1521)) based on the execution of the gesture tracker (1573). For example, the head-wearing electronic device (200) can identify the pose and / or gesture of the user's hand based on data (or images) acquired using an external camera based on the execution of the gesture tracker (1573). The gesture tracker (1573) may be referred to as a hand tracking (HaT) module (or hand tracking program) and / or a gesture tracking module.

[0195] For example, the head-wearing electronic device (200) can identify (or track) the movement of the user's eyes of the head-wearing electronic device (200) based on the execution of the eye tracker (1574). For example, the head-wearing electronic device (200) can identify the movement of the user's eyes using data obtained from a gaze tracking camera (e.g., image sensor (1521)) based on the execution of the eye tracker (1574). The eye tracker (1574) may be referred to as an eye tracking (ET) module (or eye tracking program) and / or a gaze tracking module.

[0196] For example, the recognition service layer (1570) of the head-wearing electronic device (200) may further include a face tracker (1575) for tracking the user's face. For example, the head-wearing electronic device (200) may identify (or track) the movement of the user's face and / or the user's facial expression based on the execution of the face tracker (1575). The head-wearing electronic device (200) may estimate the user's facial expression based on the movement of the user's face based on the execution of the face tracker (1575). For example, the head-wearing electronic device (200) may identify the movement of the user's face and / or the user's facial expression based on data (e.g., images and / or videos) acquired using a camera (1525) (e.g., a camera facing at least a part of the user's face) based on the execution of the face tracker (1575).

[0197] Referring to FIG. 15, the renderer (1590) may include instructions for rendering images in a three-dimensional virtual space. A processor (1510) that executes the renderer (1590) may obtain at least one image to be displayed at least partially in a display area of ​​the display (1350) in a software application. For example, the processor (1510) that executes the renderer (1590) may determine the location of the area where an application (e.g., XR application (1542), application (1545)) will be rendered. The processor (1510) that executes the renderer (1590) may generate an image of said application to be displayed on the display (1350). The renderer (1590) may synthesize images to generate a composite image to be displayed on the display (1350).

[0198] For example, a processor (1510) that executes a renderer (1590) can divide the display area of ​​a display (1350) into a foveated portion (or may be referred to as a foveated area) and a peripheral portion (or may be referred to as a residual area) using a gaze position calculated using a position tracker (1571) and / or a gaze tracker (1574). For example, a processor (1510) that detects coordinate values ​​of the gaze position can determine the portion of the display area containing said coordinate values ​​as the foveated area. A DPU that executes a renderer (1590) can acquire at least one image corresponding to each of said foveated area and said residual area, having a size smaller than the size of the entire display area of ​​the display (1350) or having a resolution less than the resolution of the display area.

[0199] A processor (1510) that executes a renderer (1590) can obtain or generate a composite image to be displayed on a display (1350) by synthesizing an image corresponding to a foveated area and an image corresponding to a surrounding area. For example, the processor (1510) can perform upscaling to enlarge the image corresponding to the surrounding area to the size of the entire display area of ​​the display (1350). On the enlarged image, the processor (1510) can combine the image corresponding to the foveated area to generate a composite image to be displayed on the display (1350). Along the boundary line of the image corresponding to the foveated area, the processor (1510) can mix the enlarged image and the image corresponding to the foveated area by applying a visual effect such as blur.

[0200] Figure 16 shows an example of a block diagram of an electronic device for displaying an image in virtual space.

[0201] In FIG. 16, an example is described in which multiple programs / instructions are executed to display an image in a virtual space. The multiple programs / instructions may all be executed on a single processor (e.g., AP) or may be executed by multiple processors (e.g., AP, GPU (graphic processing unit), NPU (neural processing unit)). The meaning of being able to be executed by multiple processors is that some programs / instructions may be executed by a first processor and other programs / instructions may be executed by a second processor different from the first processor.

[0202] Referring to FIG. 16, a head-worn electronic device (200) may execute a virtual space manager (1650) (e.g., the virtual space manager (1551) of FIG. 15, CPM) to render an image in a virtual space. For the virtual space manager (1650), at least some of the descriptions of the virtual space manager (1551) of FIG. 15 may be referenced. The virtual space manager (1650) may include a platform for supporting virtual space services. The virtual space manager (1650) may include a runtime service (1651) (e.g., open XR runtime), a panel renderer (1652) (e.g., 2D panel render), and an XR compositor (1653) (XR compositor). Based on the execution of the runtime service (1651), the head-worn electronic device (200) may execute at least one of a user pose prediction function, a frame timing function, and / or a spatial input function. For the runtime service (1651), at least some of the descriptions of the runtime service (1552) of FIG. 15 may be referenced. The head-wearing electronic device (200) may display at least one image (video) on a panel (e.g., a 2D panel) to enable the implementation of a virtual space through a display based on the execution of panel rendering (1652). For example, the head-wearing electronic device (200) may display a rendering image corresponding to RGB information (1666) for a panel from the spatialization manager (1640) described below through a display (e.g., a display (1350)). The head-wearing electronic device (200) may composite an image of a real area (hereinafter, a pass-through image) captured through a camera in virtual space with an image of a virtual area based on the execution of an XR compositor (1653) (XR compositor). For example, the head-worn electronic device (200) can generate a composite image by merging the pass-through image and the virtual region image based on the execution of the XR synthesis unit (1653).The head-wearing electronic device (200) can transmit the generated composite image to a display buffer so that the composite image is displayed. The head-wearing electronic device (200) can identify a virtual space through a virtual space manager (1650) and display at least a portion of the virtual space on a display (1350). The virtual space manager (1650) may be referred to as a CPM. The head-wearing electronic device (200) can execute the virtual space manager (1650) to render an image corresponding to at least a portion of the virtual space.

[0203] According to one embodiment, a head-worn electronic device (200) may execute a spatialization manager (1640). The spatialization manager (1640) may perform processing for displaying an image in a three-dimensional virtual space. The head-worn electronic device (200) may perform preprocessing based on the execution of the spatialization manager (1640) so that an image can be rendered in a three-dimensional virtual space through a virtual space manager (1650). For example, the head-worn electronic device (200) may perform at least some of the functions of the renderer (1590) of FIG. 15 based on the execution of the spatialization manager (1640). The head-worn electronic device (200) may process image information provided by an application (e.g., an XR application (1610), an application providing a standard 2D screen that is not XR (1620), an application providing a system UI (1630)) based on the execution of the spatialization manager (1640). A spatialization manager (1640) (e.g., space flinger) may include a system screen manager (1641) (e.g., system scene), an input manager (1642) (e.g., input routing), and a lightweight rendering engine (1643) (e.g., impress engine). The system screen manager (1641) may be executed to display a system UI (1630). System UI-related information (1664) may be transmitted to the system screen manager (1641) from a program (e.g., API) that provides the system UI (1630). System UI-related information (1664) may be obtained through a spatializer API and / or a same-process private API. The spatialization manager (1640) may determine the layout (e.g., position, display order) of the system UI (1630) screen in three-dimensional space through pre-allocated resources.The system screen manager (1641) may transmit image information (1667) to the virtual space manager (1650) for rendering a screen of the system UI (1630) according to the layout. The input manager (1642) may be configured to process user input (e.g., user input on a system screen or app screen). The Impress engine (1643) may be a renderer for image generation (e.g., a lightweight renderer (1643)). For example, the Impress engine (1643) may be used to display the system UI (1630). According to one embodiment, the spatialization manager (1640) may include a lightweight rendering engine (1643) for rendering the system UI. According to one embodiment, if the lightweight rendering engine (1643) does not have sufficient resources to render an avatar used in an HMD, at least one external rendering engine may be used. At this time, to resolve compatibility issues with external rendering (e.g., 3rd party engine), an external rendering engine support module may be added inside the spatialization manager (1640).

[0204] According to one embodiment, the electronic device may execute an application. For example, in response to the execution of an XR application (1610) (e.g., XR application (1610), 3D game, XR map, other immersive application), the virtual space manager (1650) may be executed. The head-worn electronic device (200) may provide dual image information (1661) provided from the XR application (1610) to the virtual space manager (1650). To display images in three-dimensional space, the dual image information (1661) may include two image information that account for binocular parallax. For example, the dual image information (1661) may include a first image information for the user's left eye and a second image information for the user's right eye to render in three-dimensional virtual space. Hereinafter, the term dual image information is used in the present disclosure to refer to image information for displaying images for both eyes in three-dimensional space. In addition to the dual image information, the above dual image information may utilize binocular image information, dual image information, dual image data, dual image, binocular image data, stereoscopic image information, 3D image information, spatial image information, spatial image data, 2D-3D conversion data, dimension conversion image data, binocular parallax image data, and / or equivalent technical terms. The head-wearing electronic device (200) can generate a composite image by merging image layers through a virtual space manager (1650). The head-wearing electronic device (200) can transmit the generated composite image to a display buffer. The composite image can be displayed on the display (1350) of the head-wearing electronic device (200).

[0205] According to one embodiment, the electronic device may execute at least one application among an XR application (1610) and other applications (1620) (e.g., a first application (1620-1), a second application (1620-2), ..., a Nth application (1620-N)). According to one embodiment, the application (1620) may be configured to output image information for displaying a two-dimensional image. In other words, the application (1620) may provide a two-dimensional image. As an example, the application (1620) may be a video application, a schedule application, or an internet browser application. Let us assume that, in response to the execution of the application (1620), image information (1662) provided from the application (1620) is provided to a virtual space manager (1650). Since the image information (1662) has only x and y coordinates within a two-dimensional plane, it may be difficult to consider the order of other applications centered on the user (i.e., distance from the user). The head-worn electronic device (200) may execute a spatialization manager (1640) to provide dual image information to a virtual space manager (1650), even when displaying an application (1620) that provides a general 2D screen. For example, based on the execution of the spatialization manager (1640), the head-worn electronic device (200) may receive application-related information (1663) from the first application (1620-1). For example, application-related information (1663) may include image information representing a two-dimensional image of the first application (1620-1) (e.g., information including RGB per pixel) and / or content information in the first application (1620-1) (e.g., characteristics of content running in the first application, type of content). Application-related information (1663) may be obtained through a spatializer API.Based on the execution of the spatialization manager (1640), the head-worn electronic device (200) can identify information regarding the location of the area to be rendered and the size of the area to be rendered (hereinafter, location information). Based on the execution of the spatialization manager (1640), the head-worn electronic device (200) can generate dual image information (1665, e.g., RGBx2) that takes into account the user's binocular parallax through the image information and the location information. Based on the execution of the spatialization manager (1640), the head-worn electronic device (200) can provide the dual image information (1665) to the virtual space manager (1650). By converting a simple two-dimensional image into dual image information (1665), the problem caused by the image information (1662) being directly transmitted to the virtual space manager (1650) can be resolved. Additionally, as at least some of the functions for displaying images in virtual space are performed by the spatialization manager (1640) instead of the virtual space manager (1650), the burden on the virtual space manager (1650) may be reduced.

[0206] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0207] As described above, the head-wearing electronic device (e.g., the head-wearing electronic device (200) of FIG. 2) may include at least one processor (e.g., at least one processor (210) of FIG. 2) comprising a processing circuit, one or more cameras (e.g., one or more cameras (240) of FIG. 2), one or more sensors (e.g., one or more sensors (250) of FIG. 2), and a memory (e.g., memory (220) of FIG. 2) comprising one or more storage media configured to store one or more programs configured to be executed individually or collectively by the at least one processor. The one or more programs may include instructions that cause the head-wearing electronic device to drive the one or more cameras. The one or more programs may include instructions that cause the head-wearing electronic device to acquire sensing data from the one or more sensors while driving the one or more cameras. The one or more programs may include instructions that cause the head-worn electronic device to determine, based on the sensing data, whether an object (object (415) in FIG. 4) within the field of view (FOV) of the one or more cameras is an object of interest. The one or more programs may include instructions that cause the head-worn electronic device to add data regarding the object to a database based on determining that the object is an object of interest. The one or more programs may include instructions that cause the head-worn electronic device to determine, based on the database, user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt.

[0208] For example, the one or more programs may include instructions that cause the head-worn electronic device to identify whether the database contains a set of data related to the object, based on determining that the object is the object of interest. The one or more programs may include instructions that cause the head-worn electronic device to add the data regarding the object to the set of data contained in the database, based on identifying that the database contains the set of data.

[0209] For example, the one or more programs may include instructions that cause the head-worn electronic device to create the set of data in the database based on identifying that the database does not contain the set of data. The one or more programs may include instructions that cause the head-worn electronic device to add the data regarding the object to the set of data contained in the database.

[0210] For example, the one or more programs may include instructions that cause the head-worn electronic device to identify whether the amount of data in the database to which the data regarding the object has been added is greater than a reference amount. The one or more programs may include instructions that cause the head-worn electronic device to determine the user preference information using the database based on identifying that the amount of the data in the database is greater than the reference amount.

[0211] For example, the one or more programs may include instructions that cause the head-worn electronic device to determine the data of the other database to be added to the database by comparing the database with another database based on identifying that the amount of the data of the database is smaller than the reference amount. The one or more programs may include instructions that cause the head-worn electronic device to add the data of the other database to the database. The one or more programs may include instructions that cause the head-worn electronic device to determine the user preference information using the database to which the data of the other database has been added.

[0212] For example, the one or more programs may include instructions that cause the head-worn electronic device to detect another object corresponding to the object through the one or more cameras. The one or more programs may include instructions that cause the head-worn electronic device to generate the prompt regarding the other object based on the detection. The one or more programs may include instructions that cause the head-worn electronic device to provide the prompt along with the user preference information to the trained model. The one or more programs may include instructions that cause the head-worn electronic device to obtain information for the other object generated by the trained model using the prompt and the user preference information.

[0213] For example, the head-worn electronic device may further include a display assembly. The one or more programs may include instructions that cause the head-worn electronic device to display the information for the other object through the display assembly. The one or more programs may include instructions that cause the head-worn electronic device to identify user input that stops displaying the information for the other object. The one or more programs may include instructions that cause the head-worn electronic device to delete the data regarding the object in the database based on the user input.

[0214] For example, the data regarding the object may include an embedding vector regarding the object. The database may include a vector database.

[0215] The method described above may be performed within a head-worn electronic device comprising one or more cameras and one or more sensors. The method may include an operation of driving the one or more cameras. The method may include an operation of acquiring sensing data from the one or more sensors while driving the one or more cameras. The method may include an operation of determining, based on the sensing data, whether an object within the field of view (FOV) of the one or more cameras is an object of interest. The method may include an operation of adding data regarding the object to a database based on determining that the object is an object of interest. The method may include an operation of determining, based on the database, user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt.

[0216] For example, the above method may include an operation of identifying whether the database contains a set of data related to the object, based on determining that the object is the object of interest. The above method may include an operation of adding the data regarding the object to the set of data contained in the database, based on identifying that the database contains the set of data.

[0217] For example, the above method may include an operation to create the set of data in the database based on identifying that the database does not contain the set of data. The above method may include an operation to add the data regarding the object to the set of data contained in the database.

[0218] For example, the above method may include an operation of identifying whether the amount of data in the database to which the data regarding the object has been added is greater than a reference amount. The above method may include an operation of determining the user preference information using the database based on identifying that the amount of the data in the database is greater than the reference amount.

[0219] For example, the above method may include an operation to determine data from another database to be added to the database by comparing the database with another database based on identifying that the amount of data in the database is smaller than the reference amount. The above method may include an operation to add the data from the other database to the database. The above method may include an operation to determine the user preference information using the database to which the data from the other database has been added.

[0220] For example, the above method may include an operation of detecting another object corresponding to the object through the one or more cameras. The above method may include an operation of generating the prompt regarding the other object based on the detection. The above method may include an operation of providing the prompt together with the user preference information to the trained model. The above method may include an operation of obtaining information for the other object generated by the trained model using the prompt and the user preference information.

[0221] For example, the head-worn electronic device may further include a display assembly. The method may include an action of displaying the information for the other object through the display assembly. The method may include an action of identifying a user input that stops displaying the information for the other object. The method may include an action of deleting the data regarding the object in the database based on the user input.

[0222] For example, the data regarding the object may include an embedding vector regarding the object. The database may include a vector database.

[0223] As described above, the non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the head-wearing electronic device to drive the one or more cameras when executed by the head-wearing electronic device having one or more cameras and one or more sensors. The one or more programs may include instructions that cause the head-wearing electronic device to acquire sensing data from the one or more sensors while driving the one or more cameras when executed by the head-wearing electronic device. Based on the sensing data, the head-wearing electronic device may include instructions that cause the head-wearing electronic device to determine whether an object within the field of view (FOV) of the one or more cameras is an object of interest. The one or more programs described above may include instructions that cause the head-wearing electronic device to add data regarding the object into a database based on determining that the object is the object of interest when executed by the head-wearing electronic device. The one or more programs described above may include instructions that cause the head-wearing electronic device to determine, based on the database, user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt when executed by the head-wearing electronic device.

[0224] For example, the one or more programs may include instructions that cause the head-wearing electronic device to identify whether the database contains a set of data related to the object, based on determining that the object is the object of interest when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to add the data regarding the object to the set of data contained in the database, based on identifying that the database contains the set of data when executed by the head-wearing electronic device.

[0225] For example, the one or more programs may include instructions that cause the head-wearing electronic device to create the set of data in the database based on identifying that the database does not contain the set of data when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to add the data regarding the object to the set of data contained in the database when executed by the head-wearing electronic device.

[0226] For example, the one or more programs may include instructions that cause the head-wearing electronic device to identify whether the amount of data in the database to which the data regarding the object has been added is greater than a reference amount when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to determine the user preference information using the database based on identifying that the amount of the data in the database is greater than the reference amount when executed by the head-wearing electronic device.

[0227] For example, the one or more programs may include instructions that cause the head-wearing electronic device to determine the data of the other database to be added to the database by comparing the database with another database, based on identifying that the amount of the data of the database is smaller than the reference amount when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to add the data of the other database to the database when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to determine the user preference information using the database to which the data of the other database has been added when executed by the head-wearing electronic device.

[0228] For example, the one or more programs may include instructions that cause the head-wearing electronic device to detect another object corresponding to the object through the one or more cameras when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to generate the prompt regarding the other object based on the detection when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to provide the prompt along with the user preference information to the trained model when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to obtain information for the other object generated by the trained model using the prompt and the user preference information when executed by the head-wearing electronic device.

[0229] For example, the head-wearing electronic device may further include a display assembly. The one or more programs may include instructions that cause the head-wearing electronic device to display the information for the other object through the display assembly when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to identify user input that stops displaying the information for the other object when executed by the head-wearing electronic device. The one or more programs may include instructions that cause the head-wearing electronic device to delete the data regarding the object in the database based on the user input when executed by the head-wearing electronic device.

[0230] For example, the data regarding the object may include an embedding vector regarding the object. The database may include a vector database.

[0231] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.

Claims

1. In a head-wearable electronic device, At least one processor including a processing circuit; One or more cameras; One or more sensors; and Memory comprising one or more programs configured to be executed individually or collectively by at least one processor, and including one or more storage media, The above one or more programs are: Driving the above one or more cameras; While driving the one or more cameras mentioned above, acquiring sensing data from the one or more sensors; Based on the above sensing data, determine whether an object within the field of view (FOV) of the one or more cameras is an object of interest; Based on determining that the above object is the object of interest, add data regarding the object to the database; and Based on the above database, to determine user preference information to be provided to a model trained with the above prompt in order to refine the response regarding the prompt, Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

2. In Claim 1, The above one or more programs are: Based on determining that the above object is the object of interest, identifying whether the database contains a set of data related to the object; and Based on identifying that the above database includes the above set of data, to add the above data regarding the object to the above set of data included in the above database, Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

3. In Claim 2, The above one or more programs are: Based on identifying that the above database does not contain the above set of data, the above set of data is created within the above database; and To add the data regarding the above object to the above set of data included in the above database, Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

4. In Claim 1, The above one or more programs are: Identifying whether the amount of data in the database to which the data regarding the above object has been added is greater than a reference amount; and Based on identifying that the amount of the data in the above database is greater than the reference amount, the user preference information is determined using the above database. Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

5. In Claim 4, The above one or more programs are: Based on identifying that the amount of the data in the database is smaller than the reference amount, the data of the other database to be added to the database is determined by comparing the database with another database; Add the data from the other database mentioned above to the database; and To determine the user preference information using the database to which the data from the other database has been added, Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

6. In Claim 1, The above one or more programs are: Detecting another object corresponding to the object through the one or more cameras mentioned above; Based on the above detection, generate the above prompt regarding the other object; Providing the above prompt to the above trained model along with the above user preference information; and To obtain information for the other object generated by the trained model using the above prompt and the above user preference information, Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

7. In Claim 6, It further includes a display assembly, The above one or more programs are: Through the above display assembly, the above information for the other object is displayed; Identifying user input that stops displaying the information for the other object; and To delete the data regarding the object within the database based on the above user input, Instructions including those that cause the above-mentioned head-worn electronic device Head-worn electronic device.

8. In Claim 1, The above data regarding the above object is, Includes an embedding vector for the above object, and The above database is, including a vector database, Head-worn electronic device.

9. A method executed in a head-wearing electronic device comprising one or more cameras and one or more sensors, wherein the method comprises: Operation of driving one or more of the above cameras; The operation of acquiring sensing data from the one or more sensors while driving the one or more cameras; An operation to determine whether an object within the field of view (FOV) of the one or more cameras is an object of interest based on the above sensing data; The operation of adding data regarding the object into a database based on determining that the object is the object of interest; and Based on the above database, the operation of determining user preference information to be provided to a model trained with the prompt to refine the response regarding the prompt, method.

10. In claim 9, the method comprises: An operation to identify whether the database contains a set of data related to the object, based on determining that the object is the object of interest; and Based on identifying that the above database includes the above set of data, the operation of adding the above data regarding the object to the above set of data included in the above database, method.

11. In claim 10, the method comprises: An operation to create the set of data within the database based on the fact that the database does not contain the set of data; and The operation of adding the data regarding the above object to the set of data included in the above database, method.

12. In claim 9, the method comprises: An operation to identify whether the amount of data in the database to which the data regarding the above object has been added is greater than a reference amount; and An operation to determine user preference information using the database based on identifying that the amount of the data in the database is greater than the reference amount, method.

13. In claim 12, the method comprises: An operation to determine the data of another database to be added to the database by comparing the database with another database, based on identifying that the amount of the data of the database is smaller than the reference amount; The operation of adding the data of the aforementioned other database to the aforementioned database; and An operation including determining the user preference information using the database to which the data of the other database has been added, method.

14. In claim 9, the method comprises: The operation of detecting another object corresponding to the object through the one or more cameras mentioned above; Based on the above detection, generate the above prompt regarding the other object; The action of providing the above prompt to the above-mentioned trained model along with the above-mentioned user preference information; and A method comprising obtaining information for the other object generated by the trained model using the above prompt and the above user preference information, method.

15. In a non-transient computer-readable storage medium storing one or more programs, When the above one or more programs are executed by a head-worn electronic device having one or more cameras and one or more sensors: Driving the above one or more cameras; While driving the one or more cameras mentioned above, acquiring sensing data from the one or more sensors; Based on the above sensing data, determine whether an object within the field of view (FOV) of the one or more cameras is an object of interest; Based on determining that the above object is the object of interest, add data regarding the object to the database; and Based on the above database, to determine user preference information to be provided to a model trained with the above prompt in order to refine the response regarding the prompt, Instructions that cause the above-mentioned head-worn electronic device, Non-transient computer-readable storage media.