Electronic device and method for providing third-person perspective content

The wearable device uses a camera, sensor, and AI model to generate third-person viewpoint content, addressing the limitations of existing technologies by providing immersive experiences through event identification and content generation.

WO2025150657A1PCT designated stage expired Publication Date: 2025-07-17SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/014671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2024-09-26
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing technologies lack the ability to effectively generate third-person viewpoint content using wearable devices, such as AR glasses, that provide enhanced user experiences by integrating real-world and virtual objects, due to limitations in event identification and content generation.

Method used

The solution involves a wearable device equipped with a camera, sensor, and processor that identifies events through video and sensing data, generates descriptions, extracts prompts, and uses a generative artificial intelligence model to create third-person viewpoint content.

Benefits of technology

This approach enables the generation of immersive third-person viewpoint content, enhancing user experience by providing life logging and integrating real-world and virtual objects, thereby overcoming limitations of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024014671_17072025_PF_FP_ABST
    Figure KR2024014671_17072025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an exemplary embodiment comprises: a processor; a memory; a camera for generating a video; and a sensor for obtaining sensing data related to a user. The electronic device identifies a valid event on the basis of video and / or sensing data, extracts a prompt for generating third-person perspective content corresponding to the event, and inputs the prompt to a generative artificial intelligence model, thereby generating the content.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for providing third-person viewpoint content

[0001] The present disclosure relates to an electronic device and method for providing content from a third-person perspective.

[0002] The electronic device may include a wearable device that is worn on the user's body. For example, the wearable device may be an electronic device that provides an augmented reality (AR) service that displays computer-generated information in conjunction with external objects in the real world to provide an enhanced user experience. For example, the wearable device may include AR glasses and / or a head-mounted device (HMD). The wearable device may include a camera for capturing video and a sensor for acquiring user-related sensing data.

[0003] The above information presents background information only to aid understanding of the present disclosure. No claim or determination is made as to whether any of the above is prior art in connection with the present disclosure.

[0004] Aspects of the present disclosure address at least the aforementioned problems and / or disadvantages and provide at least the advantages described below. Accordingly, aspects of the present disclosure provide an electronic device including an antenna module.

[0005] Additional aspects will be set forth in part in the description that follows, in part will be apparent from the description, or may be learned by practice of the disclosed embodiments.

[0006] An electronic device according to one aspect of the present disclosure is provided. The electronic device may include a processor comprising processing circuitry. The electronic device may include a memory storing instructions. The electronic device may include a camera for generating a video. The electronic device may include a sensor for obtaining sensed data related to a user of the electronic device. The electronic device may include a microphone for generating audio. The instructions, when executed by the processor, may cause the electronic device to identify an event based on at least one of the video or the sensed data. The instructions, when executed by the processor, may cause the electronic device to generate a description representing the event. The instructions, when executed by the processor, may cause the electronic device to extract a prompt for generating third-person viewpoint content corresponding to the event from the description. The above instructions, when executed by the processor, may cause the electronic device to obtain the content by inputting the prompt into a generative artificial intelligence model.

[0007] A method performed by an electronic device according to another aspect of the present disclosure is provided. The method may include an operation of identifying an event based on at least one of video or sensor data. The method may include an operation of generating a description representing the event. The method may include an operation of extracting a prompt for generating third-person viewpoint content corresponding to the event from the description. The method may include an operation of obtaining the third-person viewpoint content by inputting the prompt into a generative artificial intelligence model.

[0008] According to another aspect of the present disclosure, a wearable device is provided. The wearable device may include a display configured to display visual information. The wearable device may include a camera configured to generate a video. The wearable device may include a sensor configured to acquire sensed data related to a user of the wearable device. The wearable device may include a memory storing instructions. The wearable device may include a processor including processing circuitry. The instructions, when executed by the processor, may cause the wearable device to acquire the video and the sensed data by switching the camera and the sensor to an activated state based on identifying that the wearable device is being worn. The instructions, when executed by the processor, may cause the wearable device to identify a first event based on the video. The instructions, when executed by the processor, may cause the wearable device to generate a first description representing video corresponding to a first segment in which the first event is identified. The instructions, when executed by the processor, may cause the wearable device to identify a second event based on the sensed data. The instructions, when executed by the processor, may cause the wearable device to generate a second description representing sensed data corresponding to a second segment in which the second event is identified. The instructions, when executed by the processor, may cause the wearable device to generate a third description representing a third event based on at least one of the first description or the second description.The instructions, when executed by the processor, may cause the wearable device to extract a prompt from the third description for generating third-person viewpoint content corresponding to the third event based on identifying that the third event corresponds to a valid event. The instructions, when executed by the processor, may cause the wearable device to generate the third-person viewpoint content by inputting the prompt into a generative artificial intelligence model.

[0009] In accordance with another aspect of the present disclosure, one or more non-transitory computer-readable storage media are provided that store computer-executable instructions that, when executed individually or collectively by processors, cause an electronic device to perform tasks. The tasks include identifying an event based on one or more of video or sensor data, generating a description representing the event, extracting a prompt for generating third-person viewpoint content corresponding to the event, and inputting the prompt into a generative artificial intelligence model to obtain the third-person viewpoint content.

[0010] Other aspects, advantages, and key features of the present disclosure will become apparent to those skilled in the art from the following detailed description of various embodiments of the present disclosure, taken in conjunction with the accompanying drawings.

[0011] The above-described and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0012] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment;

[0013] FIG. 2 is a block diagram illustrating components of an electronic device according to an exemplary embodiment;

[0014] FIGS. 3A and 3B are flow charts illustrating operations of an electronic device generating third-person viewpoint content according to an exemplary embodiment;

[0015] Figure 4 is a flowchart showing operations of an electronic device performed by execution of a first application;

[0016] Figure 5 illustrates an exemplary first event;

[0017] Figure 6 illustrates an exemplary second event;

[0018] Figure 7 is a flowchart showing operations of an electronic device performed by execution of a second application;

[0019] Figure 8 illustrates an example of content from a third-person perspective;

[0020] FIG. 9a illustrates an exemplary screen representing user input for displaying a list of videos via a display;

[0021] FIG. 9b illustrates an exemplary screen in which a list of videos is displayed through a display;

[0022] FIG. 9c illustrates an exemplary screen in which content from a first-person perspective is displayed through a display within the first mode;

[0023] FIG. 9d illustrates an exemplary screen in which third-person view content is displayed through the display within the second mode;

[0024] FIG. 10 is a flowchart illustrating the operation of a first application for acquiring video or sensing data;

[0025] Figure 11 is a flowchart showing the operation of a second application that generates content;

[0026] FIG. 12A illustrates a perspective view of a wearable device according to an exemplary embodiment;

[0027] FIG. 12B illustrates one or more hardware elements arranged within a wearable device according to an exemplary embodiment; and

[0028] Figures 13a and 13b illustrate the appearance of a wearable device according to an exemplary embodiment.

[0029] It should be noted that throughout the drawings, similar reference numerals are used to describe identical or similar components, features, and structures.

[0030] The following description, with reference to the attached drawings, is provided to facilitate a comprehensive understanding of various embodiments of the present disclosure defined by the claims and their equivalents. While various specific details are included to facilitate understanding, they are to be considered merely exemplary. Accordingly, those skilled in the art will recognize that various modifications and variations can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. Furthermore, descriptions of known functions and configurations may be omitted for clarity and conciseness.

[0031] The terms and words used in the following description and claims are not intended to be limited in their bibliographic meanings, but are used solely by the inventors to facilitate a clear and consistent understanding of the present disclosure. Accordingly, it will be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only, and is not intended to limit the present disclosure, as defined by the appended claims and their equivalents.

[0032] The singular forms "a," "an," and "the" should be understood to include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more of those surfaces.

[0033] It should be understood that the blocks and combinations of flowcharts in each flowchart can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory device, or the one or more computer programs may be divided into different portions stored in different memory devices.

[0034] Any function or operation described herein may be processed by a single processor or a combination of processors. A single processor or a combination of processors is a circuit that performs processing, and includes an application processor (AP, eg, a central processing unit (CPU)), a communication processor (CP, eg, a modem), a graphics processing unit (eg, a GPU), a neural processing unit (NPU) (eg, an artificial intelligence (AI) chip), a wireless-fidelity (Wi-Fi) chip, a Bluetooth® chip, a global positioning system (GPS) chip, a near field communication (NFC) chip, connectivity chips, a sensor controller, a touch controller, a finger-print sensor controller, a display drive integrated circuit (DDI), an audio CODEC chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on chip (SoC), an integrated circuit (IC), or a similar circuit.

[0035] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.

[0036] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an external electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with an external electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with an external electronic device (104) via a server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0037] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0038] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0039] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0040] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0041] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0042] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0043] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0044] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., an external electronic device (102)) (e.g., a speaker or headphones) directly or wirelessly connected to the electronic device (101).

[0045] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0046] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) to an external electronic device (e.g., the external electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0047] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., an external electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0048] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0049] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0050] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0051] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0052] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., external electronic device (102), external electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0053] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., an external electronic device (104)), or a network system (e.g., a second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0054] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0055] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0056] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0057] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0058] According to an example embodiment, the electronic device (101) may include a wearable device. For example, the electronic device (101) may include a head-mounted display (HMD) that is wearable on a user's head. The electronic device (101) may be referred to as a head-mounted display (HMD) device, a headgear electronic device (101), glasses-type (or goggle-type) electronic device (101), a video see-through (VST) device, an extended reality (XR) device, a virtual reality (VR) device, and / or an augmented reality (AR) device. For example, the electronic device (101) may include an accessory (e.g., a strap) for attaching to a user's head. An example of a hardware configuration included in the electronic device (101) is described below with reference to FIG. 2. The wearable device (1200) illustrated in FIGS. 12A and 12B has an external appearance in the form of glasses, but the embodiment is not limited thereto. An example of the structure of a wearable device that can be worn on a user's head is described below with reference to FIGS. 13A and 13B.

[0059] An electronic device (101) according to an exemplary embodiment may perform functions related to augmented reality (AR) and / or mixed reality (MR). For example, when a user wears the electronic device (101), the electronic device (101) may include at least one lens positioned adjacent to the user's eyes. The electronic device (101) may combine ambient light passing through the lens with light emitted from a display (e.g., the display (250) of FIG. 2). A display area of ​​the display (250) may be formed within the lens through which the ambient light passes. Since the electronic device (101) combines the ambient light and the light emitted from the display (250), the user may see an image in which a real object recognized by the ambient light and a virtual object formed by the light emitted from the display (250) are mixed. The augmented reality, mixed reality, and / or virtual reality described above may be referred to as extended reality (XR). The electronic device (101) according to an exemplary embodiment may perform functions related to video see-through (VST) and / or virtual reality (VR).

[0060] FIG. 2 is a block diagram illustrating components of an electronic device according to an exemplary embodiment.

[0061] Referring to FIG. 2, an electronic device (101) according to an exemplary embodiment may include a camera (230) (e.g., the camera module of FIG. 1), a sensor (240) (e.g., the sensor module (176) of FIG. 1), a memory (220) (e.g., the memory (130) of FIG. 1), a processor (210) (e.g., the processor (120) of FIG. 1), a communication circuit (260) (e.g., the communication module (190) of FIG. 1), and / or a microphone (270) (e.g., the audio module (170) of FIG. 1). The microphone (270), the camera (230), the sensor (240), the memory (220), the processor (210), and / or the communication circuit (260) may be electronically and / or operably coupled with each other by an electronic component, such as a communication bus (201). The type and / or number of hardware components included in the electronic device (101) are not limited to those illustrated in FIG. 2. For example, the electronic device (101) may include only some of the hardware components illustrated in FIG. 2, or may include hardware components not illustrated in FIG. 2 (e.g., the battery module of FIG. 1).

[0062] According to an exemplary embodiment, the processor (210) may control the operation of the electronic device (101). The processor (210) may include a hardware component for processing data based on instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a field programmable gate array (FPGA), a central processing unit (CPU), and / or an application processor (AP). In an exemplary embodiment, the electronic device (101) may include one or more processors. The processor (210) may have a multi-core processor structure such as a dual core, a quad core, a hexa core, and / or an octa core. The multi-core processor structure of the processor (210) may include a structure based on a plurality of core circuits (e.g., a big-little structure) that are distinguished by power consumption, clock, and / or calculation amount per unit time. In embodiments including a processor having a multi-core processor architecture, the operations and / or functions of the present disclosure may be collectively performed by one or more cores included in the processor (210).

[0063] According to an exemplary embodiment, the memory (220) may include a hardware component for storing data and / or instructions input and / or output to the processor (210). The memory (220) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, and embedded multi media card (eMMC). In one embodiment, the memory (220) may be referred to as storage.

[0064] According to an exemplary embodiment, the display (250) can output visual information to a user of the electronic device (101). When the user wears the electronic device (101), the display (250) can be arranged in front of the user's eyes. For example, the display (250) can be controlled by a processor (210) including a circuit such as a graphic processing unit (GPU) to output visualized information to the user. The display (250) can include a flexible display (250), a flat panel display (FPD), and / or electronic paper. The display (250) can include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs can include organic LEDs (OLEDs). The embodiment is not limited thereto, and for example, if the electronic device (101) includes a lens for transmitting external light (or ambient light), the display (250) may include a projector (or projection assembly) for projecting light onto the lens. The display (250) may also be referred to as a display panel and / or a display (250) module. When a user wears the electronic device (101), pixels included in the display (250) may be arranged to face one of the user's two eyes. For example, the display (250) may include display areas (or active areas) corresponding to each of the user's two eyes.

[0065] According to an exemplary embodiment, the camera (230) may include one or more optical sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate electrical signals representing the color and / or brightness of light. The camera (230) may also be referred to as an image sensor. The plurality of optical sensors included in the camera (230) may be arranged in the form of a two-dimensional array. The camera (230) may acquire electrical signals of each of the plurality of optical sensors substantially simultaneously and generate two-dimensional frame data corresponding to light reaching the optical sensors of the two-dimensional array. For example, photographic data captured using the camera (230) may mean one (a) two-dimensional frame data acquired from the camera (230). For example, video data captured using the camera (230) may mean a sequence of a plurality of two-dimensional frame data acquired from the camera (230) according to a frame rate. The camera (230) may further include a flash light that is positioned toward the direction in which the camera (230) receives light and outputs light toward the direction.

[0066] For example, the camera (230) may be positioned toward the external environment of a user wearing the electronic device (101). The camera (230) may be positioned toward the external environment within the electronic device (101) to capture the external environment. The processor (210) may identify one or more objects using images and / or videos acquired from the camera (230). For example, the processor (210) may be configured to identify one or more objects located within the external environment based on images and / or videos acquired from the camera (230).

[0067] According to one embodiment, the microphone (270) may be configured to convert a sound wave, which is an analog signal received from outside the electronic device (101), into an electrical signal (e.g., an audio signal). For example, the microphone (270) may include a diaphragm configured to generate an electrical signal by vibrating based on the sound wave. For example, the microphone (270) may also be implemented as a part of the camera (230). For example, a video generated by the camera (230) may include audio captured by the microphone (270) in addition to an image.

[0068] According to an exemplary embodiment, the electronic device (101) may include a plurality of cameras positioned toward different directions. For example, the electronic device (101) may include a gaze tracking camera. The gaze tracking camera may be positioned toward at least one of the two eyes of a user wearing the electronic device (101). The processor (210) may identify the direction of the user's gaze using images and / or videos acquired from the gaze tracking camera. The gaze tracking camera (230) may include an infrared (IR) sensor. The gaze tracking camera (230) may also be referred to as an eye sensor, a gaze tracker, and / or an eye tracker.

[0069] According to an exemplary embodiment, the sensor (240) may generate electrical information that may be processed and / or stored by the processor (210) from non-electronic information related to the electronic device (101) and / or a user of the electronic device (101). The information may be referred to as sensing data. The sensor (240) may include a global positioning system (GPS) sensor for detecting a geographic location of the electronic device (101), an audio sensor (e.g., a microphone and / or a microphone array including multiple microphones), an ambient light sensor, an inertial measurement unit (IMU) (e.g., an acceleration sensor, a gyro sensor, and / or a geomagnetic sensor), and / or a time-of-flight (ToF) sensor (or a ToF camera). According to an exemplary embodiment, the electronic device (101) may include a multi-modal sensor.

[0070] According to an exemplary embodiment, the communication circuit (260) may include circuitry for supporting transmission and / or reception of electrical signals between the electronic device (101) and an external electronic device. The communication circuit (260) may include, for example, at least one of a modem (MODEM), an antenna, and an optical / electronic (O / E) converter. The communication circuit (260) may support transmission and / or reception of electrical signals based on various types of protocols, such as Ethernet, a local area network (LAN), a wide area network (WAN), wireless fidelity (WiFi), Bluetooth, Bluetooth low energy (BLE), ZigBee, long term evolution (LTE), 5G new radio (NR), 6G and / or above-6G. In one embodiment, the communication circuit (260) may be referred to as a communication processor (210) and / or a communication module.

[0071] According to an exemplary embodiment, instructions representing data to be processed, calculations to be performed, and / or operations to be performed by the processor (210) may be stored within the memory (220). A set of instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and / or a software application (hereinafter, “application”). For example, the electronic device (101) and / or the processor (210) may perform at least one of the operations of FIGS. 3A, 3B, 4, and 7 when a set of a plurality of instructions distributed in the form of an operating system, firmware, driver, program, and / or application is executed. That an application is installed in an electronic device (101) may mean that instructions provided in the form of an application are stored in a memory (220), and that the applications are stored in a format executable by the processor (210) (e.g., a file with an extension specified by the operating system of the electronic device (101)). For example, the application may include a program and / or a library related to a service provided to a user.

[0072] Referring to FIG. 2, programs installed in an electronic device (101) may be included in any one of different layers, including a framework layer (220a) and an application layer (220b), based on the target. The layers illustrated in FIG. 2 are logically (or for convenience of explanation) separated and may not mean that the address space of the memory (220) is separated by the layers.

[0073] According to an exemplary embodiment, within the framework layer (220a), programs (e.g., a first application (221) and / or a second application (222)) designed to target at least one of the hardware (e.g., a camera (230), a sensor (240), a memory (220), a processor (210), and / or a communication circuit (260)) of the electronic device (101) and / or the application layer (220b) may be included. The programs included in the framework layer (220a) may provide an API (application programming interface) that is executable (or callable) based on another program.

[0074] According to an exemplary embodiment, the first application (221) may be used to detect an event and a section in which an event occurred based on video and / or sensing data. The first application (221) may be operatively connected to a camera (230) and / or a sensor (240). For example, the first application (221) may include a data acquisition unit (221a), a data storage unit (221b), an action recognition unit (221c), and / or an event detection unit (221d). The data acquisition unit (221a) may be instructions or code for acquiring video captured by the camera (230) and / or sensing data generated by the sensor (240). The data storage unit (221b) may be instructions or code for storing video and / or sensing data. The action recognition unit (221c) may be instructions or codes for detecting a user's action based on video and / or sensing data. The event detection unit (221d) may be instructions or codes for detecting the occurrence of an event based on video and / or sensing data. The data acquisition unit (221a), the data storage unit (221b), the action recognition unit (221c), and / or the event detection unit (221d) may be a set of instructions or codes, and may be instructions / codes or a storage space that stores instructions / codes that are at least temporarily resided in the processor (210), or may be part of the circuitry that constitutes the processor (210).

[0075] According to an exemplary embodiment, the second application (222) may be used to generate third-person viewpoint content based on data generated or processed by the first application (221). Within the present disclosure, the content may include images and videos. The second application (222) may be operatively connected to the first application (221). For example, the second application (222) may include a first-person event interpretation unit (222a), a third-person event interpretation unit (222b), a comprehensive scene interpretation unit (222c), a prompt extraction unit (222d), and / or a content generation unit (222e). The first-person event interpretation unit (222a) may be instructions or code for generating a first-person viewpoint description that interprets a detected event from a first-person viewpoint. The third-person event interpretation unit (222b) may be instructions or codes for generating a third-person viewpoint description that interprets a detected event from a third-person viewpoint. The comprehensive scene interpretation unit (222c) may be instructions or codes for generating a description (e.g., a third description) that comprehensively interprets an event. The prompt extraction unit (222d) may be instructions or codes for extracting (or generating) a prompt that can generate third-person viewpoint content by being input into a generative artificial intelligence model. The content generation unit (222e) may be instructions or codes for generating third-person viewpoint content by inputting the prompt extracted by the prompt extraction unit (222d) into a generative artificial intelligence model.The first-person event interpretation unit (222a), the third-person event interpretation unit (222b), the comprehensive scene interpretation unit (222c), the prompt extraction unit (222d), and / or the content generation unit (222e) may be a set of instructions or codes, and may be instructions / codes or a storage space storing instructions / codes that are at least temporarily resided in the processor (210), or may be part of the circuitry constituting the processor (210).

[0076] For example, a program designed to target users of the electronic device (101) may be included in the application layer (220b). Programs included in the application layer (220b) (e.g., a third application (223)) may call an application programming interface (API) to cause the execution of functions supported by programs classified into the framework layer (220a). For example, the third application (223) may include an event management screen and operation control unit (223a), an event content playback unit (223b), an event content editing unit (223c), and / or an event content input / output unit (223d) for managing, playing, editing, inputting, and / or outputting content detected as events within the XR (extended reality) system. For example, the third application (223) may display one or more visual objects on the display (250) for performing interaction with the user based on the execution of the XR system UI. A visual object may refer to an object that can be deployed on a screen for transmitting and / or interacting with information, such as text, images, icons, videos, buttons, checkboxes, text boxes, sliders, and / or tables. A visual object may be referred to as a visual guide, a virtual object, a visual element, a UI element, a view object, and / or a view element. A wearable device may provide a user with functions available within a virtual space based on the execution of an XR system UI.In the above description, the first application (221) and the second application (222) are described as being included in the framework layer (220a), and the third application (223) is described as being included in the application layer (220b), but this is not limited thereto. According to one embodiment, at least one of the first application (221), the second application (222), or the third application (223) may operate in integration with another application. In addition, at least one of the first application (221), the second application (222), or the third application (223) may not be limited to a structure in which it operates in isolation from the framework layer (220a) or the application layer (220b). According to one embodiment, the first application (221) and the second application (222) may be operations processed in hardware. For example, the first application (221) can be operated by a real-time processor, and the second application (222) can be operated by a post processor.

[0077] When the electronic device (101) according to an exemplary embodiment is implemented as a wearable device such as an HMD, a video captured by the camera (230) may include the external environment of the electronic device (101) (or the user of the electronic device (101). Since the camera (230) captures the external environment toward which the user's gaze is directed, the video may include a part of the user's body (e.g., a hand), but may not include the user's overall appearance (e.g., a face). The electronic device (101) according to an exemplary embodiment may be configured to detect an event based on video and / or sensing data and provide third-person view content representing the detected event. The third-person view content may include an avatar (e.g., avatar (810) of FIG. 8) corresponding to the user of the electronic device (101) within the content. The face of the user of the electronic device (101) may be reconstructed using previously captured content or acquired using a separate operation for collecting the user's face. Below, an exemplary operation of an electronic device (101) for providing content from a third-person viewpoint is described.

[0078] FIGS. 3A and 3B are flow charts illustrating operations of an electronic device according to an exemplary embodiment to generate content from a third-person viewpoint.

[0079] The processor (210) of FIG. 2 can perform operations of an electronic device (e.g., the electronic device (101) of FIG. 2) described with reference to FIGS. 3A and 3B. Instructions stored in a memory (e.g., the memory (220) of FIG. 2) can, when executed by the processor (210), cause the electronic device (101) to perform operations of the electronic device (101) described in FIG. 3A.

[0080] Referring to FIG. 3A, in operation 301, the processor (210) may be configured to identify a first event based on a video captured by a camera (e.g., camera (230) of FIG. 2).

[0081] In one embodiment, the first event may be referred to as an event detected through analysis of a video. In an exemplary embodiment, the camera (230) may be configured to capture a video. For example, the camera (230) may be configured to capture an external environment of the electronic device (101). The camera (230) may be positioned within the electronic device (101) in a direction toward which the user's gaze is directed. When the electronic device (101) is worn on the user's head, if the user turns or moves their head to shift their gaze, the angle of view of the camera (230) may shift, thereby changing the view of the video. In one embodiment, the captured video may be stored within the memory (220). The processor (210) may identify the first event based on the video stored within the memory (220). For example, the processor (210) may identify one or more objects within the video and identify the first event based on the one or more objects. An example of the first event is described below with reference to Fig. 5.

[0082] In one embodiment, a video may include both visual and audio information. Since the video includes audio, information regarding user speech received via a microphone (e.g., the input module (150) of FIG. 1 ) may be included within the video. The processor (210) may also identify a first event using audio (e.g., user speech) included within the video.

[0083] In operation 302, the processor (210) may be configured to generate a first description representing a video corresponding to a first segment in which a first event is identified.

[0084] According to one embodiment, the camera (230) can continuously capture video while it is activated. The captured video can be stored in the memory (220). The camera (230) can generate a video for the entire period from the start of capture to the end of capture. The video for the entire period can be stored as a single file or as multiple files divided at regular time intervals. The processor (210) can generate a first description indicating a video corresponding to a first period in which a first event is identified among the entire period. The first description is a description that interprets the video corresponding to the first period from a first-person perspective and / or a third-person perspective, and can include a description from a first-person perspective and / or a third-person perspective. The video corresponding to the first period can be referred to as a portion of the video corresponding to a first time among the entire period. For example, if the total time interval of a video is 100 minutes and the first interval is from minute 30 to minute 40 within the 100 minutes, the video corresponding to the first interval can be referred to as a video having only a timeline from minute 30 to minute 40. In one embodiment, the first interval can be distinguished based on time or by frame number.

[0085] According to an exemplary embodiment, the processor (210) may generate a first description by analyzing a video corresponding to the first segment. For example, the processor (210) may generate a first description that describes a scene of the video by analyzing the video corresponding to the first segment. For example, the processor (210) may generate the first description using an artificial intelligence model. The artificial intelligence model may be configured to detect one or more objects within the video (e.g., object detection), recognize motion of each of one or more objects (e.g., action recognition), and / or track motion of each of one or more objects (e.g., motion tracking). However, the present invention is not limited thereto.

[0086] In operation 303, the processor (210) may be configured to identify a second event based on sensing data generated by a sensor (e.g., sensor (240) of FIG. 2).

[0087] Within the present disclosure, a second event may be referred to as an event detected through analysis of sensing data. According to an exemplary embodiment, the sensor (240) may be configured to generate sensing data related to a user of the electronic device (101). For example, the sensor (240) may be a multi-modal sensor capable of acquiring various information related to the user. For example, the sensor (240) may generate data related to a user's action. The processor (210) may identify the second event by estimating the user's action based on the action-related data stored in the memory (220). For example, if the multi-modal sensor includes an image sensor, a video may be generated through the image sensor. An example of the second event is described below with reference to FIG. 6 .

[0088] In operation 304, the processor (210) may be configured to generate a second description representing sensing data corresponding to the second interval in which the second event is identified.

[0089] According to an exemplary embodiment, the sensor (240) may continuously generate sensing data while being maintained in an activated state. The generated sensing data may be stored in the memory (220). The sensor (240) may generate sensing data for the entire period from the time the sensor (240) is activated to the time it is deactivated. The processor (210) may generate a second description representing sensing data corresponding to a second period in which a second event is identified for the entire period. The second description may include a description from a first-person perspective, as a description interpreting the sensing data corresponding to the second period. The sensing data corresponding to the second period may be referred to as a portion of the sensing data corresponding to the second time among the entire period.

[0090] According to an exemplary embodiment, the processor (210) may generate a second explanation by analyzing the sensing data corresponding to the second section. For example, the processor (210) may estimate a user's action by analyzing the sensing data corresponding to the second section and generate a second explanation that explains the estimated action. For example, the processor (210) may generate the second explanation using an artificial intelligence model. The artificial intelligence model may estimate the user's motion estimated within the sensing data. However, the present invention is not limited thereto. For example, the processor (210) may also generate the second explanation by comparing the sensing data with reference data corresponding to a specific action of the user.

[0091] In operation 305, the processor (210) may be configured to generate a third description indicating a third event based on at least one of the first description or the second description.

[0092] Within the present disclosure, a third event may be referred to as an event identified as being associated with a user based on at least one of the first event or the second event. According to exemplary embodiments, the first and second descriptions may represent the same event or different events. For example, a case where the first and second descriptions represent the same event may be referred to as a case where the processor (210) detects an event from both video and sensed data, such as when a user participates in a sporting event while wearing the electronic device (101). However, the present invention is not limited thereto, and the first and second events may be independent of each other. For example, if another person appears in front of the user while the user is sitting still, the first event may be detected by an object corresponding to the other person included in the video, but the second event may not be detected because the user is sitting still. For example, if a user runs on a treadmill, a second event may be detected by the user's running action, but the video may not detect the second event because it shows substantially the same view.

[0093] According to an exemplary embodiment, the processor (210) may be configured to generate a third description indicating a third event related to the user based on at least one of the first description or the second description. Referring to the examples described above, if the user participates in a tennis match, both the first event and the second event may be detected, and thus a third description (e.g., "I am participating in a tennis match") for describing a third event corresponding to participation in a tennis match may be generated based on the first description and the second description. For example, if the user is sitting still in a room when another person opens a door and enters, the first event may be detected, and thus a third description (e.g., "Another person is opening the door and entering the room") for describing a third event corresponding to the appearance of the other person may be generated based on the first description. For example, if the user runs on a treadmill, the second event may be detected, and thus a third description (e.g., "I am running on a treadmill") for describing a third event corresponding to a running workout may be generated based on the second description.

[0094] Although the operations described in FIG. 3A are described in the order of operations 302, 303, 304, and 305, the operations of the electronic device (101) are not limited to the above order. For example, operations 302 and 303 may be performed sequentially, operations 304 and 305 may be performed sequentially, and operations 302, 303, 304, and 305 may be performed independently.

[0095] At operation 306, the processor (210) may be configured to extract a prompt for generating third-person view content corresponding to the third event based on identifying that the third event corresponds to a valid event.

[0096] According to an exemplary embodiment, the processor (210) may be configured to identify whether the third event is a valid event. An invalid event may be an event that frequently occurs in daily life, such as when a user unintentionally performs a simple action (e.g., stretching), but the first and / or second events are detected. Referring to the examples described above, the events of participating in a tennis match and running on a treadmill may be valid events, while the event of another person entering the room may not be a valid event. The identification of a valid event may be based on whether a pre-stored condition is satisfied, or may be identified through analysis using an artificial intelligence model.

[0097] In an exemplary embodiment, the processor (210) may be configured to extract a prompt for generating third-person viewpoint content corresponding to the third event, if the third event corresponds to a valid event. In an exemplary embodiment, the prompt may be referred to as a prompt that, when input to a generative artificial intelligence model, enables the output of third-person viewpoint content. To extract the prompt, the processor (210) may be configured to generate a third-person viewpoint description as well as a first-person viewpoint description when generating a third description representing the third event. For example, the third description may include a third-person viewpoint description, such as "I am playing tennis with A" or "I am running on a treadmill in gym B." Based on the third description, the processor (210) may extract a prompt for generating third-person viewpoint content for the third event.

[0098] In operation 307, the processor (210) may be configured to obtain third-person viewpoint content by inputting a prompt into a generative artificial intelligence model.

[0099] According to one embodiment, the processor (210) may be configured to obtain content corresponding to a third event by inputting a prompt into a generative artificial intelligence model. The generative artificial intelligence model may generate content from user input using an unstructured deep learning model. When a prompt is input into the generative artificial intelligence model, the prompt is converted into tokens through a text encoder, and content may be generated by denoising randomly generated noise based on the tokens. As a prompt capable of generating content from a third-person perspective is input into the generative artificial intelligence model, content from a third-person perspective may be generated. The generated content from a third-person perspective may be stored in the memory (220).

[0100] According to one embodiment, the electronic device (101) can provide content using content (e.g., thumbnails) generated by a generative artificial intelligence model and information (e.g., object information or images or users included in a video) acquired by an analyzed video or sensor. For example, the electronic device (101) can generate a prompt including relevant object information so that the generative artificial intelligence model can generate content, or control the generation of content through a deep learning model by providing a data set that can be utilized by the artificial intelligence (e.g., generative AI).

[0101] According to an exemplary embodiment, the content may include a thumbnail. For example, the thumbnail is content that summarizes and / or represents a video corresponding to a third event, and videos stored in the memory (220) may be displayed as thumbnails corresponding to each of the videos. Through the thumbnail, the user can intuitively recognize what kind of content the stored content is. According to an exemplary embodiment, if the electronic device (101) is implemented as a wearable device such as an HMD, a video captured while the user is wearing the electronic device (101) may not substantially include the user. Since content from a third-person perspective can be generated through the above-described operations, the electronic device (101) can provide life logging to the user. Through the third-person perspective content for the section where the event is detected, the user can intuitively recognize the event. In the above description, the content has been described as being provided using a generative artificial intelligence model, but is not limited thereto. For example, the electronic device (101) may acquire content using an artificial intelligence model other than a generative artificial intelligence model.

[0102] In one embodiment, while the first and second events for generating content are described as separate events, the first and second events may be a single event. For example, images and audio may be processed simultaneously via a multimodal sensor. The electronic device (101) may identify the event based on data acquired via the multimodal sensor.

[0103] Referring to FIG. 3B, at operation 311, the processor (210) may be configured to identify an event based on sensing data.

[0104] In one embodiment, an event may be identified based on sensing data acquired through a sensor (e.g., sensor (240) of FIG. 2). For example, sensor (240) may include a multi-modal sensor. A multi-modal sensor may be configured to collect information about a single object or environment in multiple modes. For example, sensor (240) may be configured to acquire image data in a visual mode and audio data in an auditory mode. Processor (210) may identify an event based on sensing data including images and audio.

[0105] In operation 313, the processor (210) may be configured to generate a description of the interval in which the event was identified.

[0106] In one embodiment, the processor (210) may analyze the sensing data to generate a description of the segment in which an event is identified. For example, the processor (210) may generate the description using an artificial intelligence model. The artificial intelligence model may analyze an event based on one or more objects or audio contained in the sensing data and generate a description representing the event.

[0107] In operation 315, the processor (210) may be configured to extract a prompt for generating third-person view content corresponding to the event.

[0108] In one embodiment, the processor (210) may extract a prompt based on a description. For example, the processor (210) may extract a prompt for generating third-person viewpoint content of an event based on a description of the event.

[0109] In operation 317, the processor (210) may be configured to obtain third-person viewpoint content by inputting a prompt into a generative artificial intelligence model.

[0110] In one embodiment, the processor (210) can obtain third-person perspective content by inputting the extracted prompt into a generative artificial intelligence model. For example, the third-person perspective content may be content that includes the user of the electronic device (101) in an event analyzed through a description of the event. As described above, the third-person perspective content can be generated from a single event identified using the sensor (240).

[0111] Figure 4 is a flowchart illustrating operations of an electronic device performed by the execution of a first application. Figure 5 illustrates an exemplary first event. Figure 6 illustrates an exemplary second event.

[0112] The processor (210) of FIG. 2 can perform operations of an electronic device (e.g., the electronic device (101) of FIG. 2) described with reference to FIG. 4. A first application (e.g., the first application (221) of FIG. 2) stored in a memory (e.g., the memory (220) of FIG. 2) can, when executed by the processor (210), cause the electronic device (101) to perform operations of the electronic device (101) described in FIG. 4.

[0113] Referring to FIG. 4, in operation 401, the processor (210) may be configured to identify wearing of the electronic device (101).

[0114] According to an exemplary embodiment, the processor (210) may be configured to identify the wearing of the electronic device (101) when the electronic device (101) is worn on the user's body. For example, the electronic device (101) may include a sensor for identifying whether the user is wearing the electronic device (101). The processor (210) may identify the wearing of the electronic device (101) based on data indicating the wearing provided from the sensor. However, the present invention is not limited thereto. For example, the processor (210) may also identify the wearing of the electronic device (101) based on a user input. For example, the processor (210) may identify the wearing of the electronic device (101) based on receiving a user input that causes the electronic device (101) to turn on.

[0115] In operation 402, the processor (210) may be configured to switch the camera (230) and sensor (240) from a deactivated state to an activated state.

[0116] According to an exemplary embodiment, the processor (210) may be configured to activate the camera (230) and the sensor (240) based on identifying the wearing of the electronic device (101). For example, the processor (210) may transmit a signal to the camera (230) and the sensor (240) to activate the camera (230) and the sensor (240) based on identifying the wearing of the electronic device (101).

[0117] In operation 403, the camera (230) can be switched to an active state.

[0118] According to an exemplary embodiment, the camera (230) may be switched from an inactive state to an active state based on receiving the signal from the processor (210).

[0119] In operation 404, the sensor (240) may be switched to an activated state.

[0120] According to an exemplary embodiment, the sensor (240) may be switched from an inactive state to an active state based on receiving the signal from the processor (210).

[0121] In operation 405, the camera (230) may be configured to generate a video.

[0122] According to an exemplary embodiment, within an activated state, the camera (230) may be configured to generate video by capturing the external environment in real time. Unless separate user input is provided, the camera (230) may remain in an activated state. The generated video may be stored in the memory (220). The video stored in the memory (220) may be referred to as a video for the entire period from the time the camera (230) is activated to the time it is deactivated. The camera (230) may generate the video using a microphone sensor. Audio may be included in the video.

[0123] In operation 406, the sensor (240) may be configured to generate sensing data.

[0124] According to an exemplary embodiment, within an activated state, the sensor (240) may be configured to generate sensing data related to the user in real time. Unless separate user input is provided, the sensor (240) may remain in the activated state. The generated sensing data may be stored in the memory (220). The sensing data stored in the memory (220) may be referred to as sensing data for the entire period from the time the sensor (240) is activated to the time it is deactivated. Logging of the sensing data may be performed after video capture by the camera (230).

[0125] According to one embodiment, a first application (e.g., the first application (221) of FIG. 2) may be executed by a third application (e.g., the third application (223) of FIG. 2). For example, the first application (221) may initialize a sensor (240) of the electronic device (101) and control the initiation of a sensing operation of the sensor (240). The first application (221) may receive a frame of a video captured through a camera (230) and store the buffer in a database within a memory (e.g., the memory (220) of FIG. 2).

[0126] In operation 407, the processor (210) may be configured to identify a first event by analyzing the motion of one or more objects within the video.

[0127] In the following description, operations 407, 408, 409, and 410 are described in that order, but this is only for convenience of description, and the operations of the electronic device (101) are not limited to the above order. For example, operations 407 and 408 may be performed sequentially, operations 409 and 410 may be performed sequentially, and operations 407 and 409 may be performed independently. As described with reference to FIG. 3B, an event may be a single event based on sensing data acquired through a multi-modal sensor. For example, the single event may be an event identified based on sensing data including an image and audio.

[0128] According to an exemplary embodiment, the processor (210) may identify a first event based on a video captured by the camera (230). For example, the processor (210) may identify one or more objects within the video and identify the first event based on the one or more objects.

[0129] Referring to FIG. 5, the first frame (501), the second frame (502), and the third frame (503) of FIG. 5 may be referred to as videos captured by the camera (230). According to an exemplary embodiment, the camera (230) may be configured to capture the external environment of the electronic device (101). The camera (230) may be positioned in the direction in which the user's (500's) gaze is directed within the electronic device (101). When the user's (500's) head is worn with the electronic device (101), if the user turns or moves his / her head to shift his / her gaze, the view of the video captured by the camera (230) may change as the angle of view of the camera (230) changes.

[0130] According to an exemplary embodiment, the first frame (501), the second frame (502), and the third frame (503) represent frames corresponding to specific sections of a video generated by the camera (230). For example, the first frame (501), the second frame (502), and the third frame (503) may be videos captured while the user (500) is sitting in a room wearing the electronic device (101). Each of the first frame (501), the second frame (502), and the third frame (503) may be frames for different times within the video. Referring to the first frame (501), objects (510), such as a door (511) and a light (512) within the room, may be included within the first frame (501). The processor (210) can detect one or more objects (510) and detect the motion of each object (510) using an artificial intelligence model capable of analyzing the video. If an event occurs in which another person (513) opens the door (511) and enters the room while the camera (230) is filming the inside of the room, the video can include a second frame (502). For example, if another person (513) opens the door (511) and enters the room, the situation in which the door (511) opens and the situation in which the other person (513) enters can be filmed, as in the second frame (502). The processor (210) can identify the occurrence of the first event by detecting the motion of the door (511) opening, the motion of the other person (513) entering the room, the motion of the other person (513) in the room, etc.

[0131] For example, if another person (513) exits the room again, the video may include a third frame (503). The processor (210) may identify the end of the first event by detecting the action of another person (513) exiting within the third frame (503). The processor (210) may identify the first interval in which the first event occurred. For example, the time information corresponding to the timing at which the first event occurred (e.g., the timing at which another person (513) opened the door (511) and entered the room) may be 00:40:00, and the time information corresponding to the timing at which the first event ended (e.g., the timing at which another person (513) exited the room) may be 00:45:00. In this case, the processor (210) may identify the time interval between 00:40:00 and 00:45:00 as the first interval.

[0132] Referring again to FIG. 4, at operation 408, the processor (210) may be configured to store the video corresponding to the first segment in the memory (220).

[0133] According to an exemplary embodiment, the processor (210) may be configured to store a video corresponding to a first period in which a first event is identified in the memory (220). For example, among the entire period, a video corresponding to a first period in which a first event is identified (e.g., a time period between 00:40:00 and 00:45:00) may be stored in the memory (220). The video corresponding to the first period may be referred to as a portion of the video for the entire period corresponding to the first period. The video corresponding to the first period may be stored in a separate storage space that is distinct from the storage space in which the video for the entire period is stored. For example, a video generated through a camera (230) may be stored in a video database in the memory (220). Within the video, a video corresponding to the first period may be stored in an event database in the memory (220). In the above description, the first period has been described as representing a specific time period of the video, but is not limited thereto. For example, information about the first segment may include time information or frame information for each frame of the video.

[0134] In operation 409, the processor (210) may be configured to identify a second event by analyzing a user's action based on the sensing data.

[0135] Figure 6 illustrates an exemplary second event.

[0136] Referring to FIG. 6, according to an exemplary embodiment, a processor (e.g., processor (210) of FIG. 2) may identify a second event based on sensing data related to a user (500) generated by a sensor (e.g., sensor (240) of FIG. 2). For example, the processor (210) may estimate an action of the user (500) based on the sensing data (611, 612, 613) and identify the second event based on the estimated action of the user (500).

[0137] The first state (601), the second state (602), and the third state (603) of FIG. 6 may be referenced as actions of the user (500). According to an exemplary embodiment, a sensor (e.g., sensor (240) of FIG. 2) may be configured to generate data (611, 612, 613) related to the user (500). According to an exemplary embodiment, the sensor (240) may include one or more sensors for obtaining data (611, 612, 613) related to the user (500). For example, the sensor (240) may include at least one of a sensor for tracking the gaze of the user (500), a sensor for obtaining data related to biometric information of the user (500), a sensor for obtaining data related to audio, or a sensor for obtaining data related to motion of the user (500). However, the present invention is not limited thereto. The biometric information of the user (500) may include information indicating the physical state and / or psychological state of the user (500), such as the body temperature, electrocardiogram, brain waves, and electrical resistance of the skin of the user (500). Data related to audio may include voice data of the user (500) and / or audio data acquired from the external environment. A sensor for acquiring data related to the motion of the user (500) may include a gyro sensor, an acceleration sensor, and / or a geomagnetic sensor for identifying the posture of the electronic device (101). A sensor for acquiring data related to the motion of the user (500) may identify the 6 degrees of freedom pose of the electronic device (101). In addition to this, various embodiments may be possible.

[0138] For example, the first state (601) may be referred to as a state in which the user (500) is not performing a separate action. The second state (602) may be referred to as a state in which the user (500) is riding a bicycle. The third state (603) may be referred to as a state in which the user (500) has finished riding a bicycle. Referring to the first state (601), while the user (500) is not performing a separate action, the sensing data (611) may represent general data. The general data may be referred to as data measured in normal times. For example, within the first state (601), data related to the motion of the user (500) may represent data corresponding to a state in which the user (500) is not performing a special action. Referring to the second state (602), while the user (500) is riding a bicycle, the sensing data (612) may represent data corresponding to an action of the user (500). According to one embodiment, the sensing data (611) may be data representing speed information, movement information, and / or location information of the user (500) using a geomagnetic sensor. Alternatively, the sensing data (611) may be data representing biometric information (e.g., heart rate, body temperature) of the user (500) using a biosensor. When the sensing data (611) represents data exceeding a threshold value or when a change in the data converges to a specific pattern, the electronic device (101) may identify the change and identify a specific action (e.g., riding a bicycle, running, etc.) of the user (500). The electronic device (101) may identify the occurrence of a second event based on identifying the specific action. When the data (611) and the data (612) are compared, the sensing data (612) acquired within the second state (620) may exhibit a more rapid change than the sensing data (611) acquired within the first state (610).For example, within the second state (602), data related to the motion of the user (500) may represent data corresponding to riding a bicycle. The processor (210) may identify the occurrence of the second event by detecting sensing data (611) corresponding to riding a bicycle.

[0139] Referring to the third state (603), if the user (500) has finished riding the bicycle, the sensing data (613) may again represent general data. The processor (210) may identify the end of the second event by detecting the sensing data that has been changed to general data. The processor (210) may identify the second section in which the second event occurred. For example, the time information corresponding to the timing at which the second event occurred (e.g., the timing at which the bicycle was started) may be 00:40:00, and the time information corresponding to the timing at which the second event ended (e.g., the timing at which the bicycle was finished) may be 00:45:00. In this case, the processor (210) may identify the time interval between 00:40:00 and 00:45:00 as the second section.

[0140] Referring again to FIG. 4, at operation 410, the processor (210) may be configured to store sensing data corresponding to the second section in the memory (220).

[0141] According to an exemplary embodiment, the processor (210) may be configured to store sensing data corresponding to a second interval in which a second event is identified in the memory (220). For example, sensing data corresponding to a second interval in which a second event is identified (e.g., a time interval between 00:40:00 and 00:45:00) among the entire interval may be stored in the memory (220). The sensing data corresponding to the second interval may be referred to as a portion corresponding to the second interval among the sensing data for the entire interval. The sensing data corresponding to the second interval may be stored in a separate storage space that is distinct from the storage space in which the sensing data for the entire interval is stored.

[0142] The aforementioned operations 404 to 410 may be performed continuously in real time while the electronic device (101) is worn. Operation 405 may be performed independently of operation 406, and operations 407 and 408 may be performed independently of operations 409 and 410. As described above, the first event identified by analyzing the video and the second event identified by analyzing the sensing data may be independent.

[0143] In operation 411, the processor (210) may be configured to identify deactivation of the electronic device (101). For example, when the user takes off the electronic device (101) or turns off the electronic device (101), the electronic device (101) may be deactivated. For example, the processor (210) may identify deactivation of the electronic device (101) based on identifying that the electronic device (101) has been separated from the user's body through a sensor for identifying whether the user is wearing the electronic device (101), or receiving a user input that causes the electronic device (101) to be turned off.

[0144] In operation 412, the processor (210) may be configured to switch the camera (230) and sensor (240) from an activated state to a deactivated state.

[0145] According to an exemplary embodiment, the processor (210) may be configured to deactivate the camera (230) and the sensor (240) based on identifying the inactivity of the electronic device (101). For example, the processor (210) may transmit a signal to the camera (230) and the sensor (240) to deactivate the camera (230) and the sensor (240) based on identifying the detachment of the electronic device (101).

[0146] In action 413, the camera (230) can be switched to an inactive state.

[0147] According to an exemplary embodiment, the camera (230) may be switched from an activated state to a deactivated state based on receiving the signal from the processor (210).

[0148] In operation 414, the sensor (240) may be switched to an inactive state.

[0149] According to an exemplary embodiment, the sensor (240) may be switched from an activated state to a deactivated state based on receiving the signal from the processor (210).

[0150] The operations described in FIG. 4 may be operations performed by the execution of a first application (221). The first application (221) may include instructions or code for storing video and sensing data, and identifying a first event and / or a second event based on the video and / or sensing data. The video corresponding to the first section and / or the sensing data corresponding to the second section stored in the memory (220) may be used to identify a third event, which is a valid event, and to obtain content from a third-person perspective. Hereinafter, operations for generating content by the execution of a second application (e.g., the second application (222) of FIG. 2) will be described.

[0151] Figure 7 is a flowchart illustrating operations of an electronic device performed by the execution of a second application. Figure 8 illustrates an example of content from a third-person perspective.

[0152] The processor (210) of FIG. 2 can perform operations of an electronic device (e.g., the electronic device (101) of FIG. 2) described with reference to FIG. 7. A second application (e.g., the second application (222) of FIG. 2) stored in a memory (e.g., the memory (220) of FIG. 2) can, when executed by the processor (210), cause the electronic device (101) to perform operations of the electronic device (101) described in FIG. 7.

[0153] Referring to FIG. 7, in operation 701, the processor (210) may be configured to generate a first description based on a video corresponding to the first section.

[0154] According to an exemplary embodiment, the processor (210) may obtain the video corresponding to the first segment, stored in operation 408 of FIG. 4, to generate content from a third-person perspective. The processor (210) may generate a first description representing the video corresponding to the first segment by analyzing the video corresponding to the first segment. The processor (210) may generate the first description by analyzing one or more objects in the video. The first description may include a description from a first-person perspective and a description from a third-person perspective. For example, a first-person event analysis unit (e.g., the first-person event analysis unit (222a) of FIG. 2) of a second application (e.g., the second application (222) of FIG. 2) may perform an interpretation of a first-person event by analyzing data. The first-person event analysis unit (222a) may generate a first-person description of the event. For example, the third-person event interpretation unit (e.g., the third-person event interpretation unit (222b) of FIG. 2) of the second application (222) can interpret a first-person event from a third-person perspective and generate a first description of the event from a third-person perspective. For example, the comprehensive scene interpretation unit (222c) of the second application (222) can generate a first description that comprehensively represents an event. For example, if a user participates in a tennis match while wearing the electronic device (101), the video can include an opposing player in the tennis match. In this case, the processor (210) can generate a first description that includes a first-person description such as “I am playing tennis with A (opponent player)” and a third-person description such as “A is playing tennis with me.” Operation 701 can be substantially the same as operation 302 of FIG. 3A. The descriptions for action 302 can be applied substantially identically to action 701.

[0155] In operation 702, the processor (210) may be configured to generate a second description based on sensing data corresponding to the second section.

[0156] According to an exemplary embodiment, the processor (210) may obtain the sensing data corresponding to the second section, stored in operation 410 of FIG. 4, to generate content from a third-person perspective. The processor (210) may analyze the sensing data corresponding to the second section to generate a second description representing the sensing data corresponding to the second section. The second description may include a description from a first-person perspective. For example, if a user participates in a tennis match while wearing the electronic device (101), the sensing data may include data related to the user's biometric information, data related to audio generated during the tennis match, and / or data related to the user's motion. Based on the data, the processor (210) may estimate an action of the user participating in the tennis match, and generate the second description based on the estimated action. For example, the processor (210) may generate the second description including a description from a first-person perspective, such as "I am participating in a tennis match." Action 702 may be substantially identical to action 304 of FIG. 3A. The descriptions for action 304 may be substantially identically applied to action 702.

[0157] In operation 703, the processor (210) may be configured to generate a third description indicating a third event based on at least one of the first description or the second description.

[0158] According to an exemplary embodiment, the processor (210) may generate a third description that synthesizes the first description and the second description. For example, since a tennis match event may cause both the first event and the second event, the first description and the second description for the tennis match may be generated. Since the third event is determined based on at least one of the first event or the second event, in this case, the third event may be determined to be a tennis match based on both the first event and the second event. The processor (210) may generate a third description, such as "I am playing tennis with A", based on the first description and the second description. Operation 703 may be substantially identical to operation 305 of FIG. 3A. The descriptions for operation 305 may be substantially identical to operation 703.

[0159] Referring to the example described above, the tennis match event may be an event that generates both the first event and the second event. For example, if a user participates in a tennis match while wearing the electronic device (101), the video and sensing data may indicate the tennis match event. Unlike the example described above, the occurrence of the third event may be determined based solely on the first event or solely on the second event. In this case, the third description may be generated based on the first description or the second description. For example, if a user views a painting while seated in a quiet room, the processor (210) may identify the occurrence of the first event by analyzing the painting included in the video. The processor (210) may generate the first description, such as "I am viewing a painting," based on the video corresponding to the first segment where the first event occurred. In the case of the painting viewing event, the second event may not be identified because sensing data may not be acquired. The processor (210) may generate a third description, such as “I am looking at a picture,” which represents a third event, which is a picture viewing event, based on the first description.

[0160] In operation 704, the processor (210) may be configured to identify whether the third event is a valid event.

[0161] According to an exemplary embodiment, the processor (210) may identify whether a third event is a valid event and determine whether to extract a third-person point of view prompt based on the identification result, in order to generate third-person point of view content for a valid event for the user. For example, when a second event is identified by the user simply stretching, a third description indicating a stretching event may be generated based on the second description. Since a stretching event is an event that may frequently occur in daily life, the processor (210) may determine that the third event is an invalid event. If the third event is identified as a valid event, operation 705 may be performed. If the third event is identified as an invalid event, it may be determined whether there is video or sensed data to be analyzed, and if there is no video or sensed data to be analyzed, the operation may be terminated.

[0162] In one embodiment, operation 704 may be optionally performed. For example, the electronic device (101) may be configured to obtain third-person viewpoint content corresponding to the third event by performing operation 705 without determining whether the third event is a valid event. For example, if the electronic device (101) always provides third-person viewpoint content for the third event, the electronic device (101) may provide third-person viewpoint content corresponding to the third event without determining whether the third event is a valid event. In this case, the electronic device (101) may be configured to provide third-person viewpoint content for all third events.

[0163] At operation 705, the processor (210) may be configured to extract a prompt from the third description.

[0164] Action 705 may be substantially identical to action 306 of FIG. 3A. The descriptions for action 306 may be substantially identically applied to action 705. For example, the processor (210) may extract a prompt for generating third-person view content of the user and A playing tennis from a third description, such as "I am playing tennis with A."

[0165] In operation 706, the processor (210) can generate third-person viewpoint content by inputting a prompt into the generative artificial intelligence model.

[0166] Action 706 may be substantially identical to action 307 of FIG. 3A. The descriptions for action 307 may be substantially identically applied to action 706.

[0167] Referring to FIG. 8, third-person view content (801) may be provided in which a user of an electronic device (101) not included in video frames (802) appears. The video frames (802) of FIG. 8 represent a plurality of frames constituting a video. In the case of a video captured by a camera (e.g., camera (230) of FIG. 2), since it is a first-person view video, the video frames (802) may only include an object (820) corresponding to an opposing player. The video frames (802) may include a part of the user's body (e.g., a hand) included in the field of view of the camera (230), but may not include the entire appearance of the user. The third-person view content (801) may include an object (820) corresponding to an opposing player and an avatar (810) corresponding to the user, and may also include an external environment (e.g., a tennis court (830)). An electronic device (101) according to an exemplary embodiment can provide life logging by providing content (801) from a third-person perspective. The electronic device (101) can provide an enhanced user experience by providing content from a third-person perspective.

[0168] According to an exemplary embodiment, an avatar (810) included in third-person view content may be preset. For example, a user may input settings related to an avatar (810) to be included in third-person view content in advance. When third-person view content is output through a generative artificial intelligence model, the content may include an avatar (810) determined based on preset user settings. However, the present invention is not limited thereto. According to an exemplary embodiment, an avatar (810) may be preset based on a plurality of contents stored in a memory (220). For example, images with a user as a subject may include an avatar (810) corresponding to the user. When the images are input to the artificial intelligence model, objects corresponding to the user included in the images may be combined to preset an avatar (810) corresponding to the user. Alternatively, the avatar (810) corresponding to the user included in the third-person content may be implemented as a default avatar according to the system settings of the electronic device (101) and may be modified through post-processing. In addition, embodiments for generating an avatar (810) corresponding to the user may vary.

[0169] According to one embodiment, the electronic device (101) may be configured to generate a third-person view thumbnail or third-person view content using first-person view content received from an external electronic device. For example, the electronic device (101) may be configured to analyze first-person view content received from an external electronic device and generate third-person view content. An avatar (810) corresponding to the user may be included within the third-person view content.

[0170] Figure 9a illustrates an exemplary screen showing user input for displaying a video list via a display. Figure 9b illustrates an exemplary screen showing a video list via a display. Figure 9c illustrates an exemplary screen showing content from a first-person perspective via a display within a first mode. Figure 9d illustrates an exemplary screen showing content from a third-person perspective via a display within a second mode.

[0171] The processor of FIG. 2 (e.g., the processor (210) of FIG. 2) can perform the operations of the electronic device (101) described with reference to FIGS. 9A, 9B, 9C, and 9D. Instructions stored in a memory (e.g., the memory (220) of FIG. 2) can, when executed by the processor (210), cause the electronic device (101) to perform the operations of the electronic device (101) illustrated in FIGS. 9A, 9B, 9C, and 9D.

[0172] Referring to FIGS. 9A, 9B, 9C, and 9D, exemplary screens (901, 902, 903, 904) displayed through a display of an electronic device (101) (e.g., display (250) of FIG. 2) are illustrated.

[0173] Referring to FIG. 9A, the electronic device (101) may display a first screen (901) through the display (250). The first screen (901) may be a screen displayed when the electronic device (101) operates in a first mode. According to an exemplary embodiment, the first mode is a mode that provides a composite image of the external environment, which may be referred to as a see-through mode or a pass-through mode. Within the first mode, the electronic device (101) may display a composite image of the external environment on the display (250). For example, in order to express an external environment existing beyond the display (250), the processor (210) may composite a virtual image with a real image acquired through the camera (230) and display the composite image on the display (250).

[0174] According to an exemplary embodiment, the electronic device (101) may display icons corresponding to each of a plurality of software applications installed in the electronic device (101) within a first screen (901). The first screen (901) is a screen for providing a list of a plurality of software applications installed in the electronic device (101) and may be referred to as a home screen and / or a launcher screen. The electronic device (101) may display a panel (920) within the first screen (901) for providing a list of frequently executed software applications. For example, the panel (920) may be referred to as a dock. Within the panel (920), the electronic device (101) may display icons (e.g., a G icon and / or an H icon) representing frequently executed software applications. The electronic device (101) may display information related to the current time and / or the battery of the electronic device (101) within the panel (920). However, the present invention is not limited thereto. According to an exemplary embodiment, icons may be displayed within the first mode. For example, the icons may be superimposed on a composite image including visual objects corresponding to objects included in the external environment (e.g., a visual object (931) corresponding to a picture of a dog and a visual object (932) corresponding to a flower pot).

[0175] According to an exemplary embodiment, the electronic device (101) may display, within the first screen (901) of FIG. 9A, an icon (910) corresponding to a software application (e.g., a gallery application) for displaying a list (e.g., a list (950) of FIG. 9B) of contents (e.g., images and / or videos) stored in the memory (220), together with the icons. While an embodiment in which the electronic device (101) displays an icon (910) representing the software application is described, embodiments are not limited thereto, and the electronic device (101) may also display text, images, and / or videos representing the software application.

[0176] According to an exemplary embodiment, the electronic device (101) may receive a first user input for executing a gallery application. The first user input may be referred to as an input for selecting an icon (910) representing the gallery application within a first screen (901). The first user input may include a hand gesture detected by a user's hand (905). For example, the electronic device (101) may obtain an image and / or video of a body part including the user's hand (905) using a camera (e.g., camera (230) of FIG. 2). Based on detecting the hand (905), the electronic device (101) may display a virtual object (940) corresponding to the hand (905) within the first screen (901). The virtual object (940) may include a three-dimensional graphical object representing a posture of the hand (905). For example, the posture and / or orientation of the virtual object (940) may correspond to the posture and / or orientation of the hand (905) detected by the electronic device (101).

[0177] According to an exemplary embodiment, while displaying a virtual object (940), the electronic device (101) may display a virtual object (941) having a shape of a line extending from the virtual object (940). The virtual object (941) may be referred to as a ray, a ray object, a cursor, a pointer, and / or a pointer object. The virtual object (941) may have a shape of a line extending from a portion of the hand (905) (e.g., a palm and / or a designated finger such as the index finger). In FIG. 9A, a virtual object (941) having a curved shape is illustrated, but the embodiment is not limited thereto. The user (110) may move the hand (905) to change the position and / or direction of the virtual object (940) and / or the virtual object (941) within the first screen (901).

[0178] According to an exemplary embodiment, while a virtual object (941) in the form of a line is displaying an exemplary first screen (901) extending toward an icon (910), the electronic device (101) may detect or identify a pinch gesture of a hand (905). For example, after acquiring an image (906) of a hand (905) in which the fingertips of all fingers included in the hand (905) are spaced apart from each other, the electronic device (101) may acquire an image (907) of a hand (905) including at least two fingers in the form of a ring, in which at least two fingertips (e.g., the thumb and the index finger) are in contact with each other. The electronic device (101) that acquires the image (907) may detect a pinch gesture expressed by at least two fingers in the form of a ring. The duration of a pinch gesture may refer to the period of time during which the fingertips of at least two fingers of a hand (905) are in contact with each other, as in image (907). The pinch gesture may correspond to, or be mapped to, a click and / or tap gesture.

[0179] According to an exemplary embodiment, in response to a pinch gesture detected while displaying the first screen (901), the electronic device (101) may launch a gallery application corresponding to the icon (910). While an exemplary operation of launching the gallery application using a hand gesture such as a pinch gesture is described, the embodiment is not limited thereto. For example, the electronic device (101) may also launch the gallery application in response to a user's utterance (e.g., "Open the gallery") and / or the pressing of a physical input button.

[0180] Referring to FIG. 9B, an exemplary second screen (902) displayed by an electronic device (101) executing a gallery application is illustrated. Referring to the second screen (902), in response to a first user input for executing the gallery application, the electronic device (101) may display, through the display (250), a list (950) of one or more videos (951, 952) stored in a memory (e.g., the memory (220) of FIG. 2). The one or more videos (951, 952) may be videos captured by a camera (e.g., the camera (230) of FIG. 2). The video captured by the camera (230) may be a first-person view video because it includes the front of the electronic device (101), and thus, an object (961) corresponding to the user (500) may be included in the video, but an avatar (962) corresponding to the user (500) of the electronic device (101) may not be included. According to an exemplary embodiment, the list (950) may be displayed in the first mode. For example, the list (950) may be superimposed with a composite image including visual objects corresponding to objects included in the external environment (e.g., a visual object (931) corresponding to a picture of a dog and a visual object (932) corresponding to a flower pot).

[0181] According to an exemplary embodiment, the list (950) may include thumbnails (951a, 951b) corresponding to each of one or more videos (951, 952). The thumbnails (951a, 951b) may be examples of the third-person viewpoint content described above. For example, the thumbnails (951a, 951b) may be third-person viewpoint content (e.g., a third-person viewpoint image or a third-person viewpoint video) representing the content of a first-person viewpoint video. For example, even if each of the one or more videos (951, 952) includes a file title, if a separate file title is not assigned, the file title may be assigned based on a specified rule (e.g., shooting time), and thus, it may be difficult for the user (500) to intuitively identify what content the one or more videos (951, 952) include through the file title. According to an exemplary embodiment, since the thumbnails (951a, 951b) are content from a third-person perspective, the thumbnails (951a, 951b) may include not only an object (961) corresponding to the other party, but also an avatar (962) corresponding to the user (500) of the electronic device (101). The user (500) can intuitively identify the video through the third-person perspective thumbnails (951a, 951b).

[0182] According to an exemplary embodiment, the electronic device (101) may receive a second user input for a thumbnail (e.g., 951a) included in the list (950). The second user input may be referred to as a user input for selecting the thumbnail (951a) within the second screen (902). The second user input may be substantially identical to the first user input. Descriptions of the first user input may be replaced with descriptions of the second user input. According to an exemplary embodiment, in response to a pinch gesture detected while displaying the second screen (902), the electronic device (101) may play a video corresponding to the thumbnail (951a).

[0183] Referring to FIG. 9C, an exemplary third screen (903) is illustrated that plays a video (981) corresponding to one thumbnail (e.g., one thumbnail (951a) of FIG. 9C). Since the video (981) is captured by a camera (230) of the electronic device (101), it may be a first-person viewpoint video.

[0184] According to an exemplary embodiment, the electronic device (101) may be configured to switch the mode of the display (250) from the first mode to the second mode based on receipt of a second user input. The video (981) may be played in the second mode, which is different from the first mode. The second mode may be referred to as a mode in which a real image is not displayed on the display (250), as opposed to a see-through mode or a pass-through mode. For example, if the user (500) turns his / her head while the video (981) is played, the view of the video (981) may change according to the movement of the user's (500) gaze. For example, when a video (981) corresponding to an event of participating in a tennis match is played, the video (981) may include an object (961) corresponding to an opposing player.

[0185] According to an exemplary embodiment, the electronic device (101) may receive a third user input for changing the viewpoint of the video (981). For example, the electronic device (101) may display a visual object (970) for the third user input for changing the viewpoint of the video (981) within the third screen (903). For example, the visual object (970) may be displayed overlapping the video (981), but is not limited thereto. For example, the visual object (970) may include text such as “Switch to third person”, but is not limited thereto. The visual object (970) may also include an icon or a button, and various embodiments are possible. Based on receiving the third user input, the electronic device (101) may perform an operation for changing the first-person viewpoint video (981) to a third-person viewpoint video (e.g., video (982) of FIG. 9D ). For example, the electronic device (101) can extract a prompt for generating the video (982) from a third-person perspective. The operation of extracting the prompt may be referred to as the operation described above. For example, the electronic device (101) can generate a video (982) changed to a third-person perspective by generating a description of the video (981), extracting a prompt based on the description, and inputting the extracted prompt into a generative artificial intelligence model.

[0186] Referring to FIG. 9d, an exemplary fourth screen (904) displayed by an electronic device (101) playing a video (982) changed to a third-person perspective is illustrated. The video (982) may be referred to as the video shown in FIG. 9c changed to a third-person perspective.

[0187] According to an exemplary embodiment, the electronic device (101) may change the mode of the display (250) from the second mode to the first mode, and play a third-person view video (982) within the first mode. The third-person view video (982) may include both an avatar (962) corresponding to the user (500) of the electronic device (101) and an object (961) corresponding to an opposing player, and may also include an external environment (e.g., a tennis court (963)). Referring to the fourth screen (904), while the third-person view video (982) is displayed within the first mode, the external environment may be displayed within the fourth screen (904). In the case of the third-person view video (982), since a change in view is not required according to a change in the direction of the user's (500) gaze, the third-person view video (982) may be played within the first mode.

[0188] When the electronic device (101) according to the exemplary embodiment is implemented as a wearable device, the mode of the display (250) may change depending on the viewpoint of the video displayed on the display (250). The electronic device (101) according to the exemplary embodiment may operate in a first mode while playing a first-person viewpoint video, and may operate in a third mode while playing a third-person viewpoint video (982). The electronic device (101) according to the exemplary embodiment may provide a video (982) changed to a third-person viewpoint, and may provide an enhanced user experience by changing the mode of the display (250) based on the viewpoint of the video. In the above-described examples, a video has been described as an example of content, but is not limited thereto. For example, the content may be an image.

[0189] Hereinafter, the entire operations of the aforementioned electronic device (101) are described based on the viewpoint that they are performed by applications (e.g., the first application (221), the second application (222), and the third application (223) of FIG. 2). The operations below may be operations performed by the applications being executed by the processor (210) of the electronic device (101).

[0190] Figure 10 is a flowchart illustrating the operation of a first application that acquires video or sensing data.

[0191] According to one embodiment, the first application (221) may be configured to store data generated during use of the electronic device (101) and detect event sections when executed by a processor (e.g., the processor (210) of FIG. 2). A user may execute the first application (221) through a third application (e.g., the third application (223) of FIG. 2).

[0192] Referring to FIG. 10, in operation 1001, the first application (221), when executed by the processor (210), may initialize and activate a sensor of the electronic device (101). In the description of FIG. 10, the sensor may be configured as a multi-modal sensor to acquire various information. For example, the sensor may be referred to as a multi-modal sensor including the aforementioned camera (e.g., camera (230) of FIG. 2), a sensor (e.g., sensor (240) of FIG. 2), and / or a microphone (e.g., microphone (270) of FIG. 2). The first application (221) may initialize the sensor so that the sensor can provide accurate information. As the sensor is activated, the sensor may acquire video and sensing data. For example, the data acquisition unit (e.g., the data acquisition unit (221a) of FIG. 2) of the first application (221) may be configured to operate a sensor and acquire video and / or sensing data generated by the sensor when executed by the processor (210).

[0193] Actions 1002, 1003, and 1004 may be referred to as actions (1000a) for video acquired by a sensor. Actions 1005, 1006, and 1007 may be referred to as actions (1000b) for sensed data acquired by a sensor and related to a user. The actions (1000a) and the actions (1000b) may be performed independently and are not limited to the order described. For example, the data storage unit of the first application (221) (e.g., the data storage unit (221b) of FIG. 2) may be configured to store video data in the database when executed by the processor (210).

[0194] In operation 1002, the first application (221), when executed by the processor (210), may receive and store a video frame. For example, the first application (221) may receive a frame of a video captured by a sensor (e.g., a camera) and store a buffer within a database. For example, the buffer may be stored within a video database of a memory (e.g., the memory (220) of FIG. 2).

[0195] In operation 1003, the first application (221), when executed by the processor (210), may analyze a video. For example, one or more objects included in the video may be identified, and the motion of the one or more identified objects may be analyzed. For example, the action recognition unit (e.g., the action recognition unit (221c) of FIG. 2) of the first application (221) may be configured to analyze a video when executed by the processor (210).

[0196] In operation 1004, the first application (221) may identify an event when executed by the processor (210). For example, the first application (221) may identify whether an event exists based on analysis of a video. If an event exists, the first application (221) may record a section in which an event is identified and store video frames for the section in an event database of the memory (220) in operation 1008. For example, the event detection unit (e.g., the event detection unit (221d) of FIG. 2) of the first application (221) may be configured to detect an event when executed by the processor (210).

[0197] In operation 1005, the first application (221) may receive and store sensing data when executed by the processor (210). For example, the first application (221) may receive sensing data acquired by a sensor (e.g., sensor (240) of FIG. 2) and store the sensing data in a sensor database of the memory (220). For example, a data storage unit (e.g., data storage unit (221b) of FIG. 2) of the first application (221) may be configured to store the sensing data in the database when executed by the processor (210).

[0198] In operation 1006, the first application (221) may analyze sensing data when executed by the processor (210). For example, the first application (221) may analyze the motion of the user of the electronic device (101) by analyzing the sensing data. As described above, the sensing data may include, but is not limited to, speed information, movement information, location information, and / or biometric information of the user. If the sensing data corresponds to reference data indicating a specific motion of the user, the first application (221) may analyze the motion of the user by identifying the specific motion of the user. For example, the action recognition unit (e.g., the action recognition unit (221c) of FIG. 2) of the first application (221) may be configured to analyze the sensing data when executed by the processor (210).

[0199] In operation 1007, the first application (221) can identify an event when executed by the processor (210). For example, the first application (221) can identify whether an event exists based on analysis of sensing data. If an event exists, the first application (221) can record a section in which an event is identified and store sensing data for the section in an event database of the memory (220) in operation 1008. For example, the event detection unit (e.g., the event detection unit (221d) of FIG. 2) of the first application (221) can be configured to detect an event when executed by the processor (210).

[0200] In operation 1009, the first application (221), when executed by the processor (210), can identify whether sensing has ended. For example, if the sensing operation of the sensor has ended, the first application (221) can end the operation. If the sensing operation of the sensor has not ended, the first application (221) can perform operations 1002 and 1005 again.

[0201] As described above, the first application (221) may be configured to store data and identify events from the data.

[0202] Figure 11 is a flowchart showing the operation of a second application that generates content.

[0203] Referring to FIG. 11, according to one embodiment, the second application (222) may be configured to detect an event section and generate content from data stored in a database when executed by a processor (e.g., the processor (210) of FIG. 2). A user may execute the second application (222) through a third application (e.g., the third application (223) of FIG. 2).

[0204] In operation 1101, the second application (222), when executed by the processor (210), may acquire an event segment from a database. For example, the second application (222) may acquire video and / or sensing data for an event segment analyzed and stored by the first application (e.g., the first application (221) of FIG. 2).

[0205] In operation 1102, the second application (222), when executed by the processor (210), may obtain video data for the event section. For example, the second application (222) may obtain video frames corresponding to the event section.

[0206] In operation 1103, the second application (222), when executed by the processor (210), can interpret a scene of the video. For example, the second application (222) can interpret the video from a first-person perspective and a third-person perspective. For example, if the video includes a screen of a conversation with another person, the interpretation from a first-person perspective can be an interpretation such as "A (another person) is having a conversation," and the interpretation from a third-person perspective can be an interpretation such as "I am having a conversation with A." ​​However, the present invention is not limited thereto. For example, the first-person event interpretation unit (e.g., the first-person event interpretation unit (222a) of FIG. 2) of the second application (222) can perform an interpretation from a first-person perspective when executed by the processor (210). For example, the third-person event interpretation unit (e.g., the third-person event interpretation unit (222b) of FIG. 2) of the second application (222) can perform interpretation of the third-person point of view when executed by the processor (210).

[0207] In operation 1104, the second application (222), when executed by the processor (210), may obtain sensing data for the event section.

[0208] In operation 1105, the second application (222) may interpret an action when executed by the processor (210). For example, the second application (222) may interpret the user's action based on sensed data for the event section. For example, if the sensed data corresponds to reference data indicating a specific action of the user, the second application (222) may interpret the specific action based on the sensed data. For example, if the sensed data corresponds to reference data indicating speed information, location information, and / or movement information corresponding to riding a bicycle, the second application (222) may interpret the user's bicycle riding action.

[0209] In operation 1106, the second application (222), when executed by the processor (210), may merge the interpretation results. For example, the second application (222) may merge the interpretation results from the video and the interpretation results from the sensing data.

[0210] In operation 1107, the second application (222) can interpret a comprehensive third-person scene when executed by the processor (210). For example, the second application (222) can perform an interpretation of the comprehensive third-person scene based on the merged interpretation results. For example, the comprehensive scene interpretation unit (e.g., the comprehensive scene interpretation unit (222c) of FIG. 2) of the second application (222) can interpret the comprehensive scene when executed by the processor (210).

[0211] In operation 1108, the second application (222), when executed by the processor (210), may identify whether the event is a valid event. For example, the second application (222) may identify whether a comprehensively interpreted event is a valid event. For example, if the comprehensively interpreted event is an event that may frequently occur in daily life (e.g., stretching), the second application (222) may identify the event as an invalid event. If the event is a valid event, operation 1109 may be performed. If the event is an invalid event, operation 1112 may be performed. According to one embodiment, operation 1108 may be omitted. For example, if the electronic device (101) is set to generate content for all events, operation 1108 may not be performed and operation 1109 may be performed.

[0212] In operation 1109, the second application (222), when executed by the processor (210), may extract a prompt for content generation. In one embodiment, the prompt may be a prompt for generating content from a third-person perspective. For example, the prompt extraction unit (e.g., the prompt extraction unit (222d) of FIG. 2) of the second application (222) may be configured to extract a prompt for content generation when executed by the processor (210).

[0213] In operation 1110, the second application (222) can generate content for the event section when executed by the processor (210). For example, the second application (222) can generate content for the event section by inputting an extracted prompt into a generative artificial intelligence model. The content can be a third-person point-of-view event video for the event section. For example, the content generation unit (e.g., the content generation unit (222e) of FIG. 2) of the second application (222) can be configured to generate content by inputting a prompt into the generative artificial intelligence model when executed by the processor (210).

[0214] In operation 1111, the second application (222) may store the event video when executed by the processor (210). For example, the second application (222) may store the generated event video in an event database of the memory.

[0215] In operation 1112, the second application (222), when executed by the processor (210), may identify whether there are any remaining analysis targets. For example, the second application (222) may identify whether there are any remaining analysis targets within the database, and if there are any remaining analysis targets, may perform operation 1101 again. If there are no remaining analysis targets, the operation may be terminated.

[0216] Hereinafter, with reference to FIGS. 12A, 12B, 13A, and / or 13B, an exemplary appearance of a wearable device is illustrated as an example of the aforementioned electronic device (101). The wearable device (1200) of FIGS. 12A and 12B and / or the wearable device (1300) of FIGS. 13A and 13B may be an example of the aforementioned electronic device (101).

[0217] FIG. 12A illustrates an example of a perspective view of a wearable device according to an exemplary embodiment. FIG. 12B illustrates one or more hardware elements arranged within a wearable device according to an exemplary embodiment.

[0218] A wearable device (1200) according to an exemplary embodiment may have the form of glasses that are wearable on a body part of a user (e.g., head). The wearable device (1200) may include a head-mounted display (HMD). For example, the housing of the wearable device (1200) may include a flexible material, such as rubber and / or silicone, that is configured to fit closely to a portion of the user's head (e.g., a portion of the face surrounding both eyes). For example, the housing of the wearable device (1200) may include one or more straps that are capable of being twined around the user's head, and / or one or more temples that are detachably attachable to the ears of the head.

[0219] Referring to FIG. 12A, a wearable device (1200) according to an exemplary embodiment may include at least one display (1250) and a frame supporting at least one display (1250).

[0220] A wearable device (1200) according to an exemplary embodiment can be worn on a part of a user's body. The wearable device (1200) can provide augmented reality (AR), virtual reality (VR), or mixed reality (MR) that combines augmented reality and virtual reality to a user wearing the wearable device (1200). For example, the wearable device (1200) can display a virtual reality image provided from at least one optical device (1282, 1284) of FIG. 12B on at least one display (1250) in response to a user's designated gesture acquired through the motion recognition cameras (1260-2, 1260-3) of FIG. 12B.

[0221] According to an exemplary embodiment, at least one display (1250) may provide visual information to a user. For example, at least one display (1250) may include a transparent or translucent lens. At least one display (1250) may include a first display (1250-1) and / or a second display (1250-2) spaced apart from the first display (1250-1). For example, the first display (1250-1) and the second display (1250-2) may be positioned at positions corresponding to the user's left and right eyes, respectively.

[0222] Referring to FIG. 12B, at least one display (1250) can provide visual information transmitted from external light to a user through a lens included in at least one display (1250), and other visual information distinct from the visual information. The lens can be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens. For example, at least one display (1250) can include a first surface (1231) and a second surface (1232) opposite to the first surface (1231). A display area can be formed on the second surface (1232) of at least one display (1250). When a user wears the wearable device (1200), external light can be transmitted to the user by being incident on the first surface (1231) and transmitted through the second surface (1232). For another example, at least one display (1250) can display an augmented reality image combined with a virtual reality image provided from at least one optical device (1282, 1284) on a real screen transmitted through external light, in a display area formed on the second surface (1232).

[0223] In an exemplary embodiment, at least one display (1250) may include at least one waveguide (1233, 1234) that diffracts light emitted from at least one optical device (1282, 1284) and transmits the diffracted light to a user. The at least one waveguide (1233, 1234) may be formed based on at least one of glass, plastic, or polymer. A nano-pattern may be formed on at least a portion of the exterior or interior of the at least one waveguide (1233, 1234). The nano-pattern may be formed based on a grating structure having a polygonal and / or curved shape. Light incident on one end of the at least one waveguide (1233, 1234) may be propagated to the other end of the at least one waveguide (1233, 1234) by the nano-pattern. At least one waveguide (1233, 1234) may include at least one diffractive element (e.g., a diffractive optical element (DOE), a holographic optical element (HOE)) and at least one reflective element (e.g., a reflective mirror). For example, at least one waveguide (1233, 1234) may be arranged within the wearable device (1200) to guide a screen displayed by at least one display (1250) to the user's eyes. For example, the screen may be transmitted to the user's eyes based on total internal reflection (TIR) ​​occurring within the at least one waveguide (1233, 1234).

[0224] The wearable device (1200) can analyze an object included in a real image collected through a shooting camera (1260-4), combine a virtual object corresponding to an object to be provided with augmented reality among the analyzed objects, and display the virtual object on at least one display (1250). The virtual object can include at least one of text and an image regarding various information related to the object included in the real image. The wearable device (1200) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the wearable device (1200) can perform spatial recognition (e.g., simultaneous localization and mapping (SLAM)) using a multi-camera and / or time-of-flight (ToF). A user wearing the wearable device (1200) can view an image displayed on at least one display (1250).

[0225] According to an exemplary embodiment, the frame may be formed as a physical structure that allows the wearable device (1200) to be worn on the user's body. According to an exemplary embodiment, the frame may be configured so that, when the user wears the wearable device (1200), the first display (1250-1) and the second display (1250-2) can be positioned corresponding to the user's left and right eyes. The frame may support at least one display (1250). For example, the frame may support the first display (1250-1) and the second display (1250-2) to be positioned corresponding to the user's left and right eyes.

[0226] Referring to FIG. 12A, the frame may include a region (1220) that at least partially contacts a portion of the user's body when the user wears the wearable device (1200). For example, the region (1220) of the frame that contacts a portion of the user's body may include a region that contacts a portion of the user's nose, a portion of the user's ear, and a portion of the side of the user's face that the wearable device (1200) makes contact with. According to an exemplary embodiment, the frame may include a nose pad (1210) that contacts a portion of the user's body. When the wearable device (1200) is worn by the user, the nose pad (1210) may contact a portion of the user's nose. The frame may include a first temple (1204) and a second temple (1205) that contact another portion of the user's body that is distinct from the portion of the user's body.

[0227] For example, the frame may include a first rim (1201) that surrounds at least a portion of the first display (1250-1), a second rim (1202) that surrounds at least a portion of the second display (1250-2), a bridge (1203) that is disposed between the first rim (1201) and the second rim (1202), a first pad (1211) that is disposed along a portion of an edge of the first rim (1201) from one end of the bridge (1203), a second pad (1212) that is disposed along a portion of an edge of the second rim (1202) from the other end of the bridge (1203), a first temple (1204) that extends from the first rim (1201) and is secured to a portion of an ear of the wearer, and a second temple (1205) that extends from the second rim (1202) and is secured to a portion of an ear opposite the ear. The first pad (1211) and the second pad (1212) may be in contact with a portion of the user's nose, and the first temple (1204) and the second temple (1205) may be in contact with a portion of the user's face and a portion of the user's ear. The temples (1204, 1205) may be rotatably connected to the rim through the hinge units (1206, 1207) of FIG. 12B. The first temple (1204) may be rotatably connected to the first rim (1201) through the first hinge unit (1206) disposed between the first rim (1201) and the first temple (1204). The second temple (1205) may be rotatably connected to the second rim (1202) via a second hinge unit (1207) disposed between the second rim (1202) and the second temple (1205). According to an exemplary embodiment, the wearable device (1200) may use a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of a surface of the frame to identify an external object (e.g., a user's fingertip) touching the frame and / or a gesture performed by the external object.

[0228] According to an exemplary embodiment, the wearable device (1200) may include hardwares (e.g., hardwares described above based on the block diagram of FIG. 2) that perform various functions. For example, the hardwares may include a battery module (1270), an antenna module (1275), at least one optical device (1282, 1284), speakers (e.g., speakers 1255-1, 1255-2), microphones (e.g., microphones 1265-1, 1265-2, 1265-3), a light-emitting module, and / or a printed circuit board (PCB) (1290). The various hardware components may be arranged within a frame.

[0229] According to an exemplary embodiment, microphones (e.g., microphones 1265-1, 1265-2, 1265-3) of the wearable device (1200) may be disposed on at least a portion of the frame to acquire sound signals. A first microphone (1265-1) disposed on the bridge (1203), a second microphone (1265-2) disposed on the second rim (1202), and a third microphone (1265-3) disposed on the first rim (1201) are illustrated in FIG. 12B , but the number and arrangement of the microphones (1265) are not limited to the exemplary embodiment of FIG. 12B . When the number of microphones (1265) included in the wearable device (1200) is two or more, the wearable device (1200) may identify the direction of the sound signal by using multiple microphones disposed on different portions of the frame.

[0230] According to an exemplary embodiment, at least one optical device (1282, 1284) may project a virtual object onto at least one display (1250) to provide various image information to a user. For example, at least one optical device (1282, 1284) may be a projector. At least one optical device (1282, 1284) may be disposed adjacent to at least one display (1250) or may be included within at least one display (1250) as a part of at least one display (1250). According to an exemplary embodiment, the wearable device (1200) may include a first optical device (1282) corresponding to a first display (1250-1) and a second optical device (1284) corresponding to a second display (1250-2). For example, at least one optical device (1282, 1284) may include a first optical device (1282) disposed at an edge of a first display (1250-1) and a second optical device (1284) disposed at an edge of a second display (1250-2). The first optical device (1282) may transmit light to a first waveguide (1233) disposed on the first display (1250-1), and the second optical device (1284) may transmit light to a second waveguide (1234) disposed on the second display (1250-2).

[0231] In an exemplary embodiment, the camera (1260) may include a recording camera (1260-4), an eye tracking camera (ET CAM) (1260-1), and / or a motion recognition camera (1260-2, 1260-3). The recording camera (1260-4), the eye tracking camera (1260-1), and the motion recognition cameras (1260-2, 1260-3) may be positioned at different locations on the frame and may perform different functions. The eye tracking camera (1260-1) may output data indicating the position or gaze of the eyes of a user wearing the wearable device (1200). For example, the wearable device (1200) may detect the gaze from an image including the user's pupils obtained through the eye tracking camera (1260-1).

[0232] The wearable device (1200) can identify an object (e.g., a real object and / or a virtual object) focused on by the user using the user's gaze acquired through the gaze tracking camera (1260-1). The wearable device (1200) that has identified the focused object can execute a function (e.g., gaze interaction) for interaction between the user and the focused object. The wearable device (1200) can express a part corresponding to the eye of an avatar representing the user in a virtual space using the user's gaze acquired through the gaze tracking camera (1260-1). The wearable device (1200) can render an image (or screen) displayed on at least one display (1250) based on the position of the user's eyes.

[0233] For example, the visual quality of a first region related to the gaze within an image and the visual quality (e.g., resolution, brightness, saturation, grayscale, PPI (pixels per inch)) of a second region distinct from the first region may be different from each other. The wearable device (1200) may obtain an image having the visual quality of the first region matching the user's gaze and the visual quality of the second region using foveated rendering. For example, if the wearable device (1200) supports an iris recognition function, user authentication may be performed based on iris information obtained using the gaze tracking camera (1260-1). Although an example in which the gaze tracking camera (1260-1) is positioned toward the user's right eye is illustrated in FIG. 12B, the embodiment is not limited thereto, and the gaze tracking camera (1260-1) may be positioned solely toward the user's left eye, or toward both eyes.

[0234] In an exemplary embodiment, the capturing camera (1260-4) can capture an actual image or background to be aligned with a virtual image to implement augmented reality or mixed reality content. The capturing camera (1260-4) can be used to obtain a high-resolution image based on HR (high resolution) or PV (photo video). The capturing camera (1260-4) can capture an image of a specific object existing at a location where the user is looking and provide the image to at least one display (1250). The at least one display (1250) can display a single image in which information about an actual image or background including an image of the specific object obtained using the capturing camera (1260-4) and a virtual image provided through at least one optical device (1282, 1284) are superimposed. The wearable device (1200) can compensate for depth information (e.g., the distance between the wearable device (1200) and an external object acquired through a depth sensor) using an image acquired through the capture camera (1260-4). The wearable device (1200) can perform object recognition using an image acquired using the capture camera (1260-4). The wearable device (1200) can perform a function of focusing on an object (or subject) in an image (e.g., auto focus) and / or an optical image stabilization (OIS) function (e.g., anti-shake function) using the capture camera (1260-4). The wearable device (1200) can perform a pass-through function to display an image acquired through the capture camera (1260-4) by overlapping at least a portion of a screen representing a virtual space on at least one display (1250) while displaying a screen. In an exemplary embodiment, the shooting camera (1260-4) may be positioned on a bridge (1203) disposed between the first rim (1201) and the second rim (1202).

[0235] The gaze tracking camera (1260-1) can implement more realistic augmented reality by tracking the gaze of a user wearing a wearable device (1200) and thereby matching the user's gaze with visual information provided to at least one display (1250). For example, when the wearable device (1200) looks straight ahead, the wearable device (1200) can naturally display environmental information related to the user's front at a location where the user is located on at least one display (1250). The gaze tracking camera (1260-1) can be configured to capture an image of the user's pupil to determine the user's gaze. For example, the gaze tracking camera (1260-1) can receive gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In an exemplary embodiment, the gaze tracking camera (1260-1) can be positioned at positions corresponding to the user's left and right eyes. For example, the gaze tracking camera (1260-1) may be positioned within the first rim (1201) and / or the second rim (1202) to face the direction in which the user wearing the wearable device (1200) is positioned.

[0236] The gesture recognition camera (1260-2, 1260-3) can recognize the movement of the user's entire body, such as the user's torso, hand, or face, or a part of the body, and thereby provide a specific event on a screen provided on at least one display (1250). The gesture recognition camera (1260-2, 1260-3) can recognize the user's gesture (gesture recognition), obtain a signal corresponding to the gesture, and provide a display corresponding to the signal on at least one display (1250). The processor can identify the signal corresponding to the gesture, and perform a designated function based on the identification. The gesture recognition camera (1260-2, 1260-3) can be used to perform a spatial recognition function using SLAM and / or a depth map for 6 degrees of freedom pose (6 dof pose). The processor may perform gesture recognition and / or object tracking functions using the motion recognition cameras (1260-2, 1260-3). In an exemplary embodiment, the motion recognition cameras (1260-2, 1260-3) may be positioned on the first rim (1201) and / or the second rim (1202).

[0237] The camera (1260) included in the wearable device (1200) is not limited to the above-described gaze tracking camera (1260-1) and motion recognition cameras (1260-2, 1260-3). For example, the wearable device (1200) can identify an external object included in the user's field of view (FoV) using a camera positioned toward the FoV. The wearable device (1200) can identify an external object based on a sensor for identifying the distance between the wearable device (1200) and the external object, such as a depth sensor and / or a time of flight (ToF) sensor. The camera (1260) positioned toward the FoV can support an autofocus function and / or an optical image stabilization (OIS) function. For example, the wearable device (1200) may include a camera (1260) (e.g., a face tracking (FT) camera) positioned toward the face to obtain an image including the face of a user wearing the wearable device (1200).

[0238] Although not shown, the wearable device (1200) according to an exemplary embodiment may further include a light source (e.g., an LED) that emits light toward a subject (e.g., a user's eyes, face, and / or an external object within the FoV) being photographed using the camera (1260). The light source may include an infrared wavelength LED. The light source may be disposed on at least one of the frame and hinge units (1206, 1207).

[0239] According to an exemplary embodiment, a battery module (1270) may supply power to electronic components of a wearable device (1200). In an exemplary embodiment, the battery module (1270) may be disposed within the first temple (1204) and / or the second temple (1205). For example, the battery module (1270) may be a plurality of battery modules (1270). The plurality of battery modules (1270) may be disposed within each of the first temple (1204) and the second temple (1205). In an exemplary embodiment, the battery module (1270) may be disposed at an end of the first temple (1204) and / or the second temple (1205).

[0240] The antenna module (1275) can transmit signals or power to the outside of the wearable device (1200), or receive signals or power from the outside. In an exemplary embodiment, the antenna module (1275) can be positioned within the first temple (1204) and / or the second temple (1205). For example, the antenna module (1275) can be positioned close to one surface of the first temple (1204) and / or the second temple (1205).

[0241] The speaker (1255) can output an acoustic signal to the outside of the wearable device (1200). The acoustic output module may be referred to as a speaker. In an exemplary embodiment, the speaker (1255) may be positioned within the first temple (1204) and / or the second temple (1205) so as to be positioned adjacent to the ear of a user wearing the wearable device (1200). For example, the speaker (1255) may include a second speaker (1255-2) positioned within the first temple (1204) and thus adjacent to the user's left ear, and a first speaker (1255-1) positioned within the second temple (1205) and thus adjacent to the user's right ear.

[0242] The light-emitting module (not shown) may include at least one light-emitting element. The light-emitting module may emit light of a color corresponding to a specific state or emit light with an action corresponding to a specific state in order to visually provide information regarding a specific state of the wearable device (1200) to the user. For example, when the wearable device (1200) requires charging, it may emit red light at a regular cycle. In an exemplary embodiment, the light-emitting module may be disposed on the first rim (1201) and / or the second rim (1202).

[0243] Referring to FIG. 12B, according to an exemplary embodiment, a wearable device (1200) may include a printed circuit board (PCB) (1290). The PCB (1290) may be included in at least one of the first temple (1204) or the second temple (1205). The PCB (1290) may include an interposer disposed between at least two sub-PCBs. One or more hardwares included in the wearable device (1200) (e.g., hardwares illustrated by different blocks in FIG. 2) may be disposed on the PCB (1290). The wearable device (1200) may include a flexible PCB (FPCB) for interconnecting the hardwares.

[0244] According to an exemplary embodiment, a wearable device (1200) may include at least one of a gyro sensor, a gravity sensor, and / or an acceleration sensor for detecting a posture of the wearable device (1200) and / or a posture of a body part (e.g., a head) of a user wearing the wearable device (1200). Each of the gravity sensor and the acceleration sensor may measure gravitational acceleration and / or acceleration based on mutually perpendicular designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). The gyro sensor may measure an angular velocity of each of the designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). At least one of the gravity sensor, the acceleration sensor, and the gyro sensor may be referred to as an inertial measurement unit (IMU). According to an exemplary embodiment, the wearable device (1200) may identify a user's motion and / or gesture performed to execute or terminate a specific function of the wearable device (1200) based on the IMU.

[0245] Figures 13a and 13b illustrate the appearance of a wearable device according to an exemplary embodiment.

[0246] The wearable device (1300) of FIGS. 13A and 13B may include at least a portion of the hardware of the wearable device (1200) described with reference to FIGS. 12A and / or 12B. An example of an appearance of a first side (1310) of a housing of the wearable device (1300) according to an exemplary embodiment may be illustrated in FIG. 13A, and an example of an appearance of a second side (1320) opposite to the first side (1310) may be illustrated in FIG. 13B.

[0247] Referring to FIG. 13A, a first surface (1310) of a wearable device (1300) according to an exemplary embodiment may have a form attachable on a body part of a user (e.g., the face of the user). Although not shown, the wearable device (1300) may further include a strap for fixing on a body part of a user, and / or one or more temples (e.g., the first temple (1204) and / or the second temple (1205) of FIGS. 12A and 12B). A first display (1250-1) for outputting an image to a left eye among the user's two eyes, and a second display (1250-2) for outputting an image to a right eye among the two eyes, may be disposed on the first surface (1310). The wearable device (1300) may be formed on the first surface (1310) and may further include a rubber or silicone packing to prevent interference from light (e.g., ambient light) different from the light emitted from the first display (1250-1) and the second display (1250-2).

[0248] According to an exemplary embodiment, a wearable device (1300) may include cameras (1260-1) for photographing and / or tracking both eyes of a user adjacent to each of the first display (1250-1) and the second display (1250-2). The cameras (1260-1) may be referred to as the gaze tracking camera (1260-1) of FIG. 12B. According to an exemplary embodiment, a wearable device (1300) may include cameras (1260-5, 1260-6) for photographing and / or recognizing a face of a user. The cameras (1260-5, 1260-6) may be referred to as FT cameras. The wearable device (1300) can control an avatar representing the user in a virtual space based on the facial motion of the user identified using cameras (1260-5, 1260-6). For example, the wearable device (1300) can change the texture and / or shape of a part of the avatar (e.g., a part of the avatar representing a human face) using information obtained by cameras (1260-5, 1260-6) (e.g., FT camera) and representing the facial expression of the user wearing the wearable device (1300).

[0249] Referring to FIG. 13B, a camera (e.g., cameras 1260-7, 1260-8, 1260-9, 1260-10, 1260-11, 1260-12)) and / or a sensor (e.g., a depth sensor 1330) may be disposed on a second surface (1320) opposite to the first surface (1310) of FIG. 13A to obtain information related to the external environment of the wearable device (1300). For example, the cameras (1260-7, 1260-8, 1260-9, 1260-10) may be disposed on the second surface (1320) to recognize external objects. The cameras (1260-7, 1260-8, 1260-9, 1260-10) of FIG. 13b can correspond to the motion recognition cameras (1260-2, 1260-3) of FIG. 12b.

[0250] For example, using cameras (1260-11, 1260-12), the wearable device (1300) can obtain images and / or videos to be transmitted to each of the user's eyes. The camera (1260-11) can be placed on the second face (1320) of the wearable device (1300) to obtain an image to be displayed through the second display (1250-2) corresponding to the right eye among the two eyes. The camera (1260-12) can be placed on the second face (1320) of the wearable device (1300) to obtain an image to be displayed through the first display (1250-1) corresponding to the left eye among the two eyes. The cameras (1260-11, 1260-12) can correspond to the shooting camera (1260-4) of FIG. 12B.

[0251] According to an exemplary embodiment, the wearable device (1300) may include a depth sensor (1330) disposed on the second face (1320) to identify a distance between the wearable device (1300) and an external object. Using the depth sensor (1330), the wearable device (1300) may obtain spatial information (e.g., a depth map) for at least a portion of the FoV of a user wearing the wearable device (1300). Although not shown, a microphone may be disposed on the second face (1320) of the wearable device (1300) to obtain a sound output from an external object. The number of microphones may be one or more depending on the embodiment.

[0252] An electronic device (101) is provided. The electronic device (101) may include a processor (210) including processing circuitry. The electronic device (101) may include a memory (220) for storing instructions. The electronic device (101) may include a camera (230) for generating a video. The electronic device (101) may include a sensor (240) for obtaining sensing data related to a user of the electronic device (101). The electronic device (101) may include a microphone (270) for generating audio. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify an event based on at least one of the video or the sensing data. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate a description representing the event. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to extract a prompt for generating third-person viewpoint content corresponding to the event from the description. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to obtain the third-person viewpoint content by inputting the prompt into a generative artificial intelligence model.

[0253] For example, the instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify a first event based on the video. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate a first description representing video corresponding to a first segment in which the first event was identified. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify a second event based on the sensed data. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate a second description representing sensed data corresponding to a second segment in which the second event was identified. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate a third description indicating a third event based on at least one of the first description or the second description. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to extract, from the third description, a prompt for generating third-person viewpoint content corresponding to the third event. The third event may be an event identified as an event related to the user based on at least one of the first event or the second event. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate the third-person viewpoint content by inputting the prompt into the generative artificial intelligence model.

[0254] For example, the instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to extract, from the third description, the prompt for generating third-person viewpoint content corresponding to the third event based on identifying that the third event corresponds to a valid event.

[0255] For example, the third-person view content may include a thumbnail corresponding to the video.

[0256] For example, the electronic device (101) may further include a display (250) for displaying visual information. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to receive a first user input for displaying a video list, including the thumbnail corresponding to the video, through the display (250). The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to display the video list through the display (250) based on receipt of the first user input. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to receive a second user input for a thumbnail within the video list. The above instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to play a video corresponding to the one thumbnail through the display (250) based on receipt of the second user input.

[0257] For example, the instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify the event based on one or more objects within the video.

[0258] For example, the instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to estimate a user action based on the sensing data. The instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify the event based on the estimated user action.

[0259] For example, the memory (220) may store a first software application (221). The first software application (221), when individually or collectively executed by the processor (210), may cause the electronic device (101) to store the video and the sensing data generated in real time. The first software application (221), when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify a first event based on the video. The first software application (221), when individually or collectively executed by the processor (210), may cause the electronic device (101) to store the video corresponding to the first section in which the first event is identified in the memory (220). The first software application (221), when individually or collectively executed by the processor (210), may cause the electronic device (101) to identify a second event based on the sensing data. The first software application (221), when individually or collectively executed by the processor (210), may cause the electronic device (101) to store the sensing data corresponding to the second section in which the second event is identified in the memory (220).

[0260] For example, the memory (220) may store a second software application (222). The second software application (222), when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate, based on the video corresponding to the first segment, the first description representing the video corresponding to the first segment. The second software application (222), when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate, based on the sensing data corresponding to the second segment, the second description representing the sensing data corresponding to the second segment.

[0261] For example, the first description may include a description from a first-person perspective and a description from a third-person perspective. The second description may include a description from a first-person perspective. The third description may include a description from a third-person perspective.

[0262] For example, the third-person view content may include an avatar corresponding to the user of the electronic device (101).

[0263] For example, the avatar corresponding to the user of the electronic device (101) may include an avatar based on an object corresponding to the user included in the third-person view content.

[0264] For example, the sensor (240) may include at least one of a sensor for tracking the user's gaze, a sensor for obtaining data related to the user's biometric information, a sensor for obtaining data related to audio, or a sensor for obtaining data related to the user's motion.

[0265] For example, the instructions, when individually or collectively executed by the processor (210), may cause the electronic device (101) to generate third-person viewpoint content corresponding to first-person viewpoint content received from an external electronic device using the generative artificial intelligence model.

[0266] For example, the electronic device (101) may include a wearable device. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to change the camera (230) and the sensor (240) from an inactive state to an active state based on the wearable device identifying the wearing of the wearable device.

[0267] For example, the electronic device (101) may include a head mounted display (HMD) device. The instructions, when individually or collectively executed by the processor (210), may cause the HMD device to receive user input for playing back the video through the display (250) in a first mode that provides a composite image of the external environment. The instructions, when individually or collectively executed by the processor (210), may cause the HMD device to change from the first mode to a second mode different from the first mode based on the receipt of the user input. The instructions, when individually or collectively executed by the processor (210), may cause the HMD device to play back the video through the display (250) in the second mode.

[0268] For example, the instructions, when individually or collectively executed by the processor (210), may cause the HMD to receive user input for changing the video to a third-person perspective while the video is being played. The instructions, when individually or collectively executed by the processor (210), may cause the HMD to extract a prompt for generating the video in a third-person perspective based on receiving the user input for changing the video to a third-person perspective. The instructions, when individually or collectively executed by the processor (210), may cause the HMD to generate the video in a third-person perspective by inputting the prompt into the generative artificial intelligence model. The instructions, when individually or collectively executed by the processor (210), may cause the HMD to change the second mode to the first mode. The above instructions, when individually or collectively executed by the processor (210), may cause the HMD to play the video from a third-person perspective within the first mode.

[0269] A method performed by an electronic device is provided. The method may include identifying an event based on video or sensor data. The method may include generating a description representing the event. The method may include extracting a prompt from the description for generating third-person viewpoint content corresponding to the event. The method may include generating the content by inputting the prompt into a generative artificial intelligence model.

[0270] For example, the content may include a thumbnail corresponding to the video.

[0271] For example, the method may further include an operation of receiving a first user input for displaying a video list including the thumbnail corresponding to the video through the display (250) of the electronic device (101). The method may further include an operation of displaying the thumbnail list through the display (250) based on the first user input. The method may further include an operation of receiving a second user input for one thumbnail included in the list. The method may further include an operation of playing a video corresponding to the one thumbnail through the display (250) based on the reception of the second user input.

[0272] For example, the method may further include an operation of identifying a first event based on one or more objects in the video. The method may further include an operation of generating a first description representing a video corresponding to a first section in which the first event is identified. The method may further include an operation of identifying a second event based on the sensed data. The method may further include an operation of generating a second description representing the sensed data corresponding to a second section in which the second event is identified. The method may further include an operation of generating a third description representing a third event based on at least one of the first description or the second description. The method may further include an operation of extracting a prompt for generating third-person viewpoint content corresponding to the third event from the third description. The method may further include an operation of generating the content by inputting the prompt into a generative artificial intelligence model.

[0273] For example, the content may include an avatar corresponding to the user.

[0274] A wearable device is provided. The wearable device may include a display (250) for displaying visual information. The wearable device may include a camera (230) for generating video. The wearable device may include a sensor (240) for obtaining sensing data related to a user of the wearable device. The wearable device may include a memory (220) for storing instructions. The wearable device may include a processor (210) including processing circuitry. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to obtain the video and the sensing data by switching the camera (230) and the sensor (240) to an active state based on identifying that the wearable device is being worn. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to identify a first event based on the video. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to generate a first description representing video corresponding to a first segment in which the first event was identified. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to identify a second event based on the sensed data. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to generate a second description representing sensed data corresponding to a second segment in which the second event was identified.The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to generate a third description indicative of a third event based on at least one of the first description or the second description. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to extract a prompt from the third description for generating third-person viewpoint content corresponding to the third event based on identifying that the third event corresponds to a valid event. The instructions, when individually or collectively executed by the processor (210), may cause the wearable device to generate the content by inputting the prompt into a generative artificial intelligence model.

[0275] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, electronic devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0276] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0277] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0278] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (120) (e.g., the processor (120)) of a machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0279] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as a memory (130) of a manufacturer's server, an application store's server, or a relay server.

[0280] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0281] It will be appreciated that the various embodiments of the present disclosure, as described in the claims and specification, may be implemented in the form of hardware, software, or a combination of hardware and software.

[0282] Such software may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores one or more computer programs (software modules), and the one or more computer programs may include computer-executable instructions that, when executed by one or more processors of the electronic device, cause the electronic device to perform the method of the present disclosure.

[0283] Any such software may be stored in the form of a volatile or non-volatile storage device, such as, for example, a read only memory (ROM), whether erasable or rewritable, or in the form of a memory, such as, for example, a random access memory (RAM), a memory chip, device, or integrated circuit, or on an optical or magnetically readable medium, such as, for example, a compact disk (CD), a digital versatile disc (DVD), a magnetic disk, or a magnetic tape. It will be appreciated that the storage device and the storage medium are various embodiments of a computer program comprising instructions that, when executed, implement various embodiments of the present disclosure, or a non-transitory machine-readable storage device suitable for storing a computer program. Accordingly, various embodiments provide a program comprising code for implementing an apparatus or method as claimed in any one of the claims herein, and a non-transitory machine-readable storage device storing such a program.

[0284] While the disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the disclosure as defined by the appended claims and their equivalents.

Claims

1. In electronic devices, A processor comprising processing circuitry; A memory comprising one or more storage media for storing instructions; A camera configured to generate video; A sensor configured to obtain sensing data related to a user of the electronic device; and comprising a microphone configured to produce audio; The above instructions, when individually or collectively executed by the processor, cause the electronic device to: Identifying an event based on at least one of the above video or the above sensing data, Generate a description representing the above event, From the above description, a prompt for generating third-person viewpoint content corresponding to the above event is extracted, By inputting the above prompt into a generative artificial intelligence model, it causes the third-person viewpoint content to be obtained. Electronic devices.

2. In paragraph 1, The above instructions, when individually or collectively executed by the processor, cause the electronic device to: Based on the above video, identify the first event, Generate a first description representing a video corresponding to the first segment in which the first event is identified, Based on the above sensing data, identify the second event, Generate a second description representing sensing data corresponding to the second segment in which the second event is identified, Generate a third description representing a third event based on at least one of the first description or the second description, From the third description, extract a prompt for generating third-person viewpoint content corresponding to the third event, wherein the third event is an event identified as an event related to the user based on at least one of the first event or the second event; By inputting the above prompt into the generative artificial intelligence model, it causes the third-person viewpoint content to be generated. Electronic devices.

3. In paragraph 2, The instructions, when individually or collectively executed by the processor, cause the electronic device to extract, from the third description, the prompt for generating third-person viewpoint content corresponding to the third event, based on identifying that the third event corresponds to a valid event. Electronic devices.

4. In any one of paragraphs 1 to 3, The above third-person view content is, Containing a thumbnail corresponding to the above video, Electronic devices.

5. In paragraph 4, Further comprising a display configured to display visual information, The above instructions, when individually or collectively executed by the processor, cause the electronic device to: Receiving a first user input for displaying a list of videos, including thumbnails corresponding to said videos, through said display; Based on receipt of the first user input, displaying the video list through the display; Receiving a second user input for a thumbnail within the above video list, Based on the reception of said second user input, causing a video corresponding to said one thumbnail to be played through said display, Electronic devices.

6. In any one of paragraphs 1 to 5, The instructions, when individually or collectively executed by the processor, cause the electronic device to identify the event based on one or more objects within the video. Electronic devices.

7. In any one of paragraphs 1 to 6, The above instructions, when individually or collectively executed by the processor, cause the electronic device to: Based on the above sensing data, the user's actions are estimated, Based on the estimated user's actions, causing the event to be identified, Electronic devices.

8. In any one of paragraphs 1 to 7, The above memory is, Save the first software application, The first software application, when executed individually or collectively by the processor, causes the electronic device to: Store the video and sensing data generated in real time, Based on the above video, identify the first event, Store the video corresponding to the first segment in which the first event is identified in the memory, Based on the above sensing data, identify the second event, Causing the sensing data corresponding to the second interval in which the second event is identified to be stored in the memory, Electronic devices.

9. In paragraph 8, The above memory is, Save the second software application, The second software application, when executed individually or collectively by the processor, causes the electronic device to: Based on the video corresponding to the first section, generate a first description representing the video corresponding to the first section, Causing to generate a second description representing the sensing data corresponding to the second section, based on the sensing data corresponding to the second section. Electronic devices.

10. In any one of paragraphs 1 to 9, The above third-person view content is, comprising an avatar corresponding to the user of the electronic device; Electronic devices.

11. In paragraph 10, The avatar corresponding to the user of the electronic device is, An avatar based on an object corresponding to the user, included within the content of the third-person viewpoint, Electronic devices.

12. In any one of paragraphs 1 to 11, The above sensor, A sensor configured to track the gaze of the user, a sensor configured to obtain data related to biometric information of the user, a sensor configured to obtain data related to audio, or a sensor configured to obtain data related to motion of the user, comprising at least one of: Electronic devices.

13. In any one of paragraphs 1 to 12, The above instructions, when individually or collectively executed by the processor, cause the electronic device to: Using the above generative artificial intelligence model, causing third-person viewpoint content to be generated corresponding to first-person viewpoint content received from an external electronic device. Electronic devices.

14. In any one of paragraphs 1 to 13, The above electronic device, Including an HMD (head mounted display) device, The above instructions, when individually or collectively executed by the processor, cause the HMD device to: In a first mode providing a synthetic image of the external environment, a user input is received for playing the video through the display, Based on the reception of the above user input, changing from the first mode to a second mode different from the first mode, In the second mode, causing the video to be played through the display, Electronic devices.

15. In paragraph 14, The above instructions, when individually or collectively executed by the processor, cause the HMD device to: While the above video is playing, user input is received to change the video to a third-person perspective, Based on the reception of the user input for changing the above video to a third-person viewpoint, extract a prompt for generating the above video in a third-person viewpoint, By inputting the above prompt into the generative artificial intelligence model, the above video is generated from a third-person perspective, Change the above second mode to the above first mode, In the above first mode, causing the video to be played from a third-person perspective, Electronic devices.

Citation Information

Patent Citations

  • Successive approximated register analog to digital converter and its opearation method

    KR1020230164450A

  • System for providing interactive content linked with automata

    KR1020230164836A

  • Systems, devices, methods and programs for controlling security using multiple artificial intelligence models

    KR102596977B1

  • Method and system for generating in-game insights

    WO2022187487A1

  • Video recording processing

    WO2023239477A1