Method and device for acquiring content related to target object

The described electronic device addresses the challenge of integrating real-world and virtual content by using sensors and eye-tracking to provide immersive and interactive augmented and mixed reality experiences.

WO2025154911A1PCT designated stage expired Publication Date: 2025-07-24SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017056
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2024-11-01
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing augmented reality and mixed reality technologies struggle to effectively integrate real-world and virtual content in a seamless manner, particularly in providing interactive and context-aware experiences.

Method used

An electronic device equipped with a display, processor, and memory is designed to obtain image information of physical objects, determine target objects, and output interactive content based on user inputs, using sensors and cameras for spatial awareness and eye-tracking to enhance user interaction.

Benefits of technology

Enables immersive and interactive augmented and mixed reality experiences by accurately overlaying virtual content onto real-world environments, enhancing user engagement and interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017056_24072025_PF_FP_ABST
    Figure KR2024017056_24072025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to the present invention includes: a display; a processor; and a memory for storing instructions. The instructions, when executed by the processor, cause the electronic device to: acquire image information corresponding to a target object selected from among physical objects arranged in a physical space in the vicinity of the electronic device; acquire content related to the target object on the basis of the acquired image information and information about the vicinity of the electronic device; deactivate an interaction object, operable in response to a user input, on the basis of satisfying a condition related to output of the content; and output the acquired content in an area corresponding to the target object through the display.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for obtaining content related to a target object

[0001] Below, a technique for obtaining content related to a target object is disclosed.

[0002] Recently, virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies that utilize computer graphics technology are being developed. VR technology uses computers to create virtual spaces that don't exist in the real world and then make them feel real, while AR or MR technologies superimpose computer-generated information on top of the real world. In other words, they combine the real and virtual worlds to enable real-time user interaction.

[0003] Among these, AR and MR technologies are being integrated and utilized in various fields, such as broadcasting, medical technology, and gaming technology. Representative examples of AR being integrated into broadcasting include the natural change of weather maps in front of a weathercaster giving a weather forecast on TV, or the insertion of non-existent advertising images into the stadium during a sports broadcast, making them appear as if they were actually there.

[0004] A representative service that provides users with augmented reality or mixed reality is the metaverse. This metaverse, a portmanteau of "meta," meaning "fictional" or "abstract," and "universe," signifies a three-dimensional virtual world. A more advanced concept than the existing term "virtual reality environment," the metaverse provides an augmented reality environment where virtual worlds like the web and the internet are integrated into the real world.

[0005] The above information may be provided as background information to aid in understanding this document. None of the above is claimed to be prior art related to this document or can be used to determine prior art.

[0006] An electronic device includes a display; at least one processor including a processing circuit; and a memory including one or more storage media storing instructions, wherein when the instructions are executed by the at least one processor, the instructions individually or collectively cause the electronic device to obtain image information corresponding to a target object determined from among physical objects arranged in a physical space around the electronic device, obtain content related to the target object based on the obtained image information and surrounding information of the electronic device, deactivate an interactive object operable in response to a user's input based on satisfying a condition regarding output of the content, and output the obtained content in an area corresponding to the target object through the display.

[0007] A method performed by an electronic device may include: an operation of acquiring image information corresponding to a target object determined from among physical objects arranged in a physical space around the electronic device; an operation of acquiring content related to the target object based on the acquired image information and surrounding information of the electronic device; an operation of deactivating an interactive object operable in response to a user's input based on satisfying a condition regarding output of the content; and an operation of outputting the acquired content in an area corresponding to the target object through a display.

[0008] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device according to various embodiments.

[0009] FIG. 2 illustrates an example of an optical see-through device according to various embodiments.

[0010] FIG. 3 illustrates examples of optical systems for an eye tracking camera, a transparent member, and a display according to various embodiments.

[0011] FIGS. 4A and 4B are drawings showing examples of the front and back of an electronic device according to various embodiments.

[0012] FIG. 5 illustrates examples of construction of a virtual space, input from a user within the virtual space, and output to the user according to various embodiments.

[0013] FIG. 6 is a diagram illustrating an example of an operation of an electronic device providing space to a user according to various embodiments.

[0014] FIG. 7 is a flowchart illustrating an example of a method for an electronic device to output content related to a target object according to various embodiments.

[0015] FIG. 8 illustrates an example of an operation of an electronic device providing a three-dimensional image of a space according to various embodiments.

[0016] FIG. 9 is a diagram illustrating an example of an operation in which the operation mode of an electronic device is switched between a basic operation mode and a rest operation mode when the target object is a movable object according to various embodiments.

[0017] FIG. 10 is a diagram illustrating an example of an operation in which the operation mode of an electronic device is switched between a basic operation mode and a rest operation mode when displaying a scene in an internal area of ​​a target object according to various embodiments.

[0018] FIG. 11 is a diagram illustrating an example of an operation of an electronic device according to various embodiments to obtain content based on a content creation model.

[0019] FIG. 12 is a diagram illustrating an example of an operation of an electronic device according to various embodiments to output content in an area larger than an area corresponding to a target object.

[0020] FIG. 13 is a diagram illustrating an example of a configuration for obtaining and / or outputting content when an electronic device selects multiple target objects according to various embodiments.

[0021] FIG. 14 is a block diagram illustrating an example configuration of an electronic device according to various embodiments.

[0022] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.

[0023] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device according to various embodiments.

[0024] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0025] The processor (120) may control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., program (140)), and may perform various data processing or operations. The processor (120) may include at least one processor including a processing circuit. According to one embodiment, as at least a part of the data processing or operation, the processor (120) may store a command or data received from another component (e.g., sensor module (176) or communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0026] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0027] The memory (130) can store various data used by at least one component (e.g., the processor (120) or the sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., the program (140)) and input data or output data for commands related thereto. The memory (130) can include a volatile memory (132) or a non-volatile memory (134). According to one embodiment, the memory (130) can include one or more storage media that store instructions that cause the operation of the electronic device (101).

[0028] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0029] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0030] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0031] A display module (160) (e.g., a display) can visually provide information to an external device (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0032] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0033] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0034] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0035] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0036] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0037] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0038] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0039] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0040] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data relation (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0041] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0042] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0043] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0044] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0045] According to one embodiment, commands or data can be transmitted or received between an electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199).

[0046] Each of the external electronic devices (102, 103) and the server (108) may be the same type of device as or different from the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 103) or the server (108). For example, when the electronic device (101) needs to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of executing the function or service itself or in addition, request one or more external electronic devices to execute at least a part of the function or service. The one or more external electronic devices that receive the request may execute at least a part of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a part of a response to the request. In this specification, an example is mainly described in which an electronic device (101) is an augmented reality device (e.g., an electronic device (201) of FIG. 2, an electronic device (301) of FIG. 3, or an electronic device (401) of FIG. 4), and an external electronic device (102, 103) or a server (108) among the servers transmits the results of executing a virtual space and additional functions or services related to the virtual space to the electronic device (101).

[0047] The server (108) may include a processor (181), a communication module (182), and a memory (183). The processor (181), the communication module (182), and the memory (183) may be configured similarly to the processor (120), the communication module (190), and the memory (130) of the electronic device (101). For example, the processor (181) may provide a virtual space and interaction between users within the virtual space by executing instructions stored in the memory (183). The processor (181) may generate at least one of visual information, auditory information, or tactile information of the virtual space and objects within the virtual space. For example, as visual information, the processor (181) may generate rendering data (e.g., visual rendering data) that renders the appearance (e.g., shape, size, color, or texture) of the virtual space and the appearance (e.g., shape, size, color, or texture) of objects positioned within the virtual space. In addition, the processor (181) may generate rendering data that renders a change (e.g., a change in the appearance of an object, a sound generation, or a tactile generation) based on at least one of an interaction between objects (e.g., a physical object, a virtual object, or an avatar object) in a virtual space, or a user's input to an object (e.g., a physical object, a virtual object, or an avatar object). The communication module (182) may establish communication with a first electronic device (e.g., an electronic device (101)) of a user and a second electronic device (e.g., an electronic device (102)) of another user. The communication module (182) may transmit at least one of the visual information, the tactile information, or the auditory information described above to the first electronic device and the second electronic device. For example, the communication module (182) may transmit rendering data.

[0048] For example, the server (108) renders content data executed in an application and transmits it to the electronic device (101), and the electronic device (101) receiving the data can output the content data to the display module (160). If the electronic device (101) detects user movement through an IMU sensor or the like, the processor (120) of the electronic device (101) can correct the rendering data received from the external electronic device (102) based on the movement information and output it to the display module (160). Alternatively, the movement information can be transmitted to the server (108) to request rendering so that the screen data is updated accordingly. However, the present invention is not limited thereto, and the aforementioned rendering can be performed by various forms of external electronic devices (102, 103), such as a smartphone or a case device capable of storing and charging the electronic device (101). Rendering data corresponding to the aforementioned virtual space generated by the external electronic device (102, 103) can be provided to the electronic device (101). For another example, the electronic device (101) may receive virtual space information (e.g., vertex coordinates, texture, color defining the virtual space) and object information (e.g., vertex coordinates, texture, color defining the appearance of an object) from the server (108) and perform rendering on its own based on the received data.

[0049] FIG. 2 illustrates an example of an optical see-through device according to various embodiments.

[0050] The electronic device (201) may include at least one of a display (e.g., a display module (160) of FIG. 1), a vision sensor, a light source (230a, 230b), an optical element, or a substrate. An electronic device (201) that has a transparent display and provides an image through the transparent display may be referred to as an optical see-through device (OST device).

[0051] The display may include, for example, a liquid crystal display (LCD), a digital mirror device (DMD), a liquid crystal on silicon (LCoS), an organic light emitting diode (OLED), or a micro light emitting diode (micro LED).

[0052] In one embodiment, when the display is formed of one of a liquid crystal display (LCD), a digital mirror display (DMD), or a silicon liquid crystal display (SLCD), the electronic device (201) may include a light source (230a, 230b) that irradiates light to a screen output area of ​​the display (e.g., a screen display unit (215a, 215b)). In another embodiment, when the display can generate light on its own, for example, when it is formed of one of an organic light emitting diode (OLED) or a micro LED, the electronic device (201) may provide a good quality virtual image to the user even without including a separate light source (230a, 230b). In one embodiment, if the display is formed of an organic light emitting diode (OLED) or a micro LED, the light source (230a, 230b) is unnecessary, and thus the electronic device (201) may be lightweight.

[0053] Referring to FIG. 2, the electronic device (201) may include a display, a first transparent member (225a) and / or a second transparent member (225b), and a user may use the electronic device (201) while wearing it on his or her face. The first transparent member (225a) and / or the second transparent member (225b) may be formed of a glass plate, a plastic plate, or a polymer, and may be manufactured to be transparent or translucent. According to one embodiment, the first transparent member (225a) may be arranged to face the user's right eye, and the second transparent member (225b) may be arranged to face the user's left eye. The display may include a first display (205) that outputs a first image (e.g., a right image) corresponding to the first transparent member (225a), and a second display (210) that outputs a second image (e.g., a left image) corresponding to the second transparent member (225b). In one embodiment, if each display is transparent, each display and transparent member may be positioned at a position facing the user's eyes to form a screen display unit (215a, 215b).

[0054] In one embodiment, light emitted from the display (205, 210) may be guided along an optical path through an input optical member (220a, 220b) into a waveguide. Light traveling within the waveguide may be guided toward a user's eyes through an output optical member (e.g., the output optical member (340) of FIG. 3). The screen display units (215a, 215b) may be determined based on the light emitted toward the user's eyes.

[0055] For example, light emitted from the display (205, 210) may be reflected in the grating region of the waveguide formed in the input optical member (220a, 220b) and the screen display portions (215a, 215b) and transmitted to the user's eyes.

[0056] The optical element may include at least one of a lens or an optical waveguide.

[0057] The lens can adjust the focus so that the screen outputted to the display can be seen by the user. The lens may include, for example, at least one of a Fresnel lens, a pancake lens, or a multi-channel lens.

[0058] An optical waveguide can transmit image light generated from a display to a user's eyes. For example, the image light may represent light emitted by a light source (230a, 230b) that passes through a screen output area of ​​the display. The optical waveguide can be made of glass, plastic, or polymer. The optical waveguide can include a nano-pattern formed on a portion of an internal or external surface, for example, a grating structure having a polygonal or curved shape. An exemplary structure of the optical waveguide is described below in FIG. 3.

[0059] The vision sensor may include at least one of a camera sensor or a depth sensor.

[0060] The first camera (265a, 265b) is a recognition camera, and may be a camera used for 3DoF, 6DoF head tracking, hand detection, hand tracking, and spatial recognition. The first camera (265a, 265b) may mainly include a GS (Global shutter) camera. Since a stereo camera is required for head tracking and spatial recognition, the first camera (265a, 265b) may include two or more GS cameras. The GS camera may have superior performance compared to the RS (Rolling shutter) camera in terms of detecting and tracking fine movements such as rapid hand movements and fingers. For example, the GS camera may have low image blur. The first camera (265a, 265b) may capture image data used for spatial recognition for 6DoF and SLAM function through depth shooting. Additionally, a user gesture recognition function can be performed based on image data captured by the first camera (265a, 265b).

[0061] The second camera (270a, 270b) is an ET (Eye Tracking) camera and can be used to capture image data for detecting and tracking the user's pupils. The second camera (270a, 270b) is described later in FIG. 3.

[0062] The third camera (245) may be a camera for photography. The third camera (245) may include a high-resolution camera for capturing HR (High Resolution) or PV (Photo Video) images. The third camera (245) may include a color camera equipped with functions for obtaining high-quality images, such as an AF function and optical image stabilization (OIS). The third camera (245) may be a GS camera or an RS camera.

[0063] The fourth camera unit (e.g., the face recognition camera (425, 426) of FIG. 4 below) is a face recognition camera, and the FT (Face Tracking) camera can be used to detect and track the user's facial expression.

[0064] A depth sensor (not shown) may be a sensor that senses information for determining the distance to an object, such as Time of Flight (TOF). TOF is a technology that measures the distance to an object using signals (e.g., near-infrared, ultrasound, laser, etc.). A depth sensor based on TOF technology can measure the time of flight of a signal by transmitting a signal and measuring the signal at a receiver.

[0065] The light source (230a, 230b) (e.g., illumination module) may include a device (e.g., light emitting diode (LED)) that irradiates light of various wavelengths. The illumination module may be attached to various locations depending on the application. As one example, a first illumination module (e.g., LED device) attached around the frame of the augmented reality glasses device may emit light to assist in gaze detection when tracking eye movements with an ET camera. The first illumination module may include, for example, an IR LED of an infrared wavelength. As another example, a second illumination module (e.g., LED device) may be attached adjacent to a camera mounted around a hinge (240a, 240b) connecting the frame and the temple, or around a bridge connecting the frames. The second illumination module may emit light to supplement the ambient brightness when the camera is taking pictures. If subject detection is not easy in a dark environment, the second lighting module may illuminate.

[0066] A substrate (235a, 235b) (e.g., a printed circuit board (PCB)) can support the aforementioned components.

[0067] A printed circuit board (PCB) may be disposed on the temple portion of the glasses. The FPCB may transmit electrical signals to each module (e.g., a camera, a display, an audio module, a sensor module) and other printed circuit boards. In one embodiment, at least one printed circuit board may include a first substrate, a second substrate, and an interposer disposed between the first substrate and the second substrate. Electrical signals may be transmitted to each module and other printed circuit boards.

[0068] Other components may include, for example, at least one of a plurality of microphones (e.g., a first microphone (250a), a second microphone (250b), a third microphone (250c)), a plurality of speakers (e.g., a first speaker (255a), a second speaker (255b)), a battery (260), an antenna, or a sensor (e.g., an acceleration sensor, a gyro sensor, a touch sensor, etc.).

[0069] FIG. 3 illustrates examples of optical systems for an eye tracking camera, a transparent member, and a display according to various embodiments.

[0070] FIG. 3 is a diagram for explaining the operation of an eye-tracking camera included in an electronic device according to one embodiment. Referring to FIG. 3, a process is illustrated in which an eye-tracking camera (310) of an electronic device (301) according to one embodiment (e.g., the first eye-tracking camera (270a) and the second eye-tracking camera (270b) of FIG. 2) tracks a user's eye (309), that is, the user's gaze, using light (e.g., infrared light) output from a display (320) (e.g., the first display (205) and the second display (210) of FIG. 2).

[0071] The second camera (e.g., the second cameras (270a, 270b) of FIG. 2) may be an eye tracking camera (310) that collects information to position the center of a virtual image projected onto the electronic device (301) according to the direction in which the pupil of the wearer of the electronic device (301) is looking. The second camera may also include a GS camera to detect the pupil and track rapid eye movement. ET cameras may also be installed for the left and right eyes, respectively, and cameras with the same performance and specifications may be used. The eye tracking camera (310) may include a gaze tracking sensor (315). The gaze tracking sensor (315) may be included inside the eye tracking camera (310). Infrared light output from the display (320) may be transmitted to the user's eyes (309) as infrared reflected light (303) by a half mirror. The gaze tracking sensor (315) can detect infrared reflected light (303) and infrared transmitted light (305) reflected from the user's eye (309). The eye tracking camera (310) can track the user's eye (309), or in other words, the user's gaze, based on the detection result of the gaze tracking sensor (315).

[0072] The display (320) may include a plurality of visible light pixels and a plurality of infrared pixels. The visible light pixels may include R, G, and B pixels. The visible light pixels may output visible light corresponding to a virtual object image. The infrared pixels may output infrared light. The display (320) may include, for example, micro light emitting diodes (LEDs) or organic light emitting diodes (OLEDs).

[0073] The display optical pipe (350) and the eye tracking camera optical pipe (360) may be included inside a transparent member (370) (e.g., the first transparent member (225a) and the second transparent member (225b) of FIG. 2). The transparent member (370) may be formed of a glass plate, a plastic plate, or a polymer, and may be manufactured to be transparent or translucent. The transparent member (370) may be positioned to face the user's eyes. At this time, the distance between the transparent member (370) and the user's eyes (309) may be referred to as 'eye relief' (380).

[0074] The transparent member (370) may include optical waveguides (350, 360). The transparent member (370) may include an input optical member (330) and an output optical member (340). In addition, the transparent member (370) may include an eye-tracking splitter (375) that separates input light into multiple waveguides.

[0075] According to one embodiment, light incident on one end of the display light pipe (350) can be propagated inside the display light pipe (350) by the nano-pattern and provided to the user. In addition, the display light pipe (350) composed of a free-form prism can provide image light to the user through the reflected light. The display light pipe (350) can include at least one diffractive element (e.g., a Diffractive Optical Element (DOE), a Holographic Optical Element (HOE)) or at least one reflective element (e.g., a reflective mirror). The display light pipe (350) can guide display light (e.g., image light) emitted from a light source to the user's eyes by using at least one diffractive element or reflective element included in the display light pipe (350). For reference, also, in FIG. 3, the output optical member (340) is expressed as being separate from the eye-tracking optical waveguide (360), but the output optical member (340) may be included inside the eye-tracking optical waveguide (360).

[0076] According to various embodiments, the diffractive element may include an input optical member (330) and an output optical member (340). For example, the input optical member (330) may mean an input grating region. The output optical member (340) may mean an output grating region. The input grating region may serve as an input terminal that diffracts (or reflects) light output from (e.g., a Micro LED) to transmit the light to a transparent member (e.g., a first transparent member, a second transparent member) of a screen display unit. The output grating region may serve as an outlet that diffracts (or reflects) light transmitted to a transparent member (e.g., a first transparent member, a second transparent member) of a waveguide to a user's eye.

[0077] According to various embodiments, the reflective element may include a total internal reflection (TIR) ​​optical element or a total internal reflection waveguide. For example, total internal reflection may refer to a method of guiding light such that light (e.g., a virtual image) entering through an input grating region is 100% reflected from one surface (e.g., a specific surface) of the waveguide, thereby transmitting 100% of the light to the output grating region.

[0078] In one embodiment, light emitted from the display (320) may be guided along an optical path through an input optical element (330) into a waveguide. Light traveling within the waveguide may be guided toward the user's eyes through an output optical element (340). The screen display may be determined based on the light emitted toward the user's eyes.

[0079] FIGS. 4A and 4B are diagrams illustrating examples of the front and back of an electronic device according to various embodiments. FIG. 4A may be an external appearance of the electronic device (401) as viewed from a first direction (①), and FIG. 4B may be an external appearance of the electronic device (401) as viewed from a second direction (②). When a user wears the electronic device (401), the external appearance viewed by the user's eyes may be FIG. 4B.

[0080] Referring to FIG. 4A, according to various embodiments, an electronic device (401) (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, and the electronic device (301) of FIG. 3) may provide a service that provides an extended reality (XR) experience to a user. For example, an XR or XR service may be defined as a service that collectively refers to virtual reality (VR), augmented reality (AR), and / or mixed reality (MR).

[0081] According to one embodiment, the electronic device (401) may refer to a head-mounted device or a head-mounted display worn on the user's head, but may also be configured in the form of at least one of glasses, goggles, a helmet, or a hat. The electronic device (401) may include an OST (optical see-through) type configured to allow external light to reach the user's eyes through the glasses when worn, or a VST (video see-through) type configured to allow light emitted from the display to reach the user's eyes but block external light so that external light does not reach the user's eyes when worn.

[0082] According to one embodiment, the electronic device (401) may be worn on the user's head and may provide the user with an image related to an extended reality (XR) service. For example, the electronic device (401) may provide XR content (hereinafter referred to as an XR content image) that outputs at least one virtual object to be superimposed on a display area or an area determined to be the user's field of view (FoV). According to one embodiment, the XR content may refer to an image or image that appears to have at least one virtual object superimposed on an image related to a real space acquired through a camera (e.g., a camera for taking pictures) or a virtual space. According to one embodiment, the electronic device (401) may provide XR content based on a function being performed by the electronic device (401) and / or a function being performed by one or more external electronic devices (e.g., the electronic devices (102, 104) of FIG. 1, the server (108) of FIG. 1).

[0083] According to one embodiment, the electronic device (401) is at least partially controlled by an external electronic device (e.g., electronic devices (102 or 104) of FIG. 1), and may perform at least one function under the control of the external electronic device, but may also perform at least one function independently.

[0084] Referring to FIG. 4A, a vision sensor may be placed on a first surface of a housing of a main body (410) of an electronic device (401). The vision sensor may include cameras (e.g., cameras for second functions (411, 412), cameras for first functions (415)) and / or a depth sensor (417) for obtaining information related to the surrounding environment of the electronic device (401).

[0085] In one embodiment, the second function cameras (411, 412) can obtain images related to the surrounding environment of the electronic device (401). The first function cameras (415) can obtain images when the wearable electronic device is worn by the user. The first function cameras (415) can be used for hand detection and tracking, and recognition of user gestures (e.g., hand movements). The first function cameras (415) can be used for 3DoF, 6DoF head tracking, position (spatial, environmental) recognition, and / or movement recognition. In one embodiment, the second function cameras (411, 412) can also be used for hand detection and tracking, and user gestures.

[0086] In one embodiment, the depth sensor (417) may be configured to transmit a signal and receive a signal reflected from a subject, and may be used for purposes such as time of flight (TOF) to determine the distance to an object. Instead of or in addition to the depth sensor (417), cameras (411, 412, 415, 416) may determine the distance to an object.

[0087] Referring to FIG. 4b, a camera (425, 426) for facial recognition and / or a display (421) (and / or lens) may be placed on the second surface (420) of the housing of the main body (410).

[0088] In one embodiment, a face recognition camera (425, 426) adjacent to the display may be used to recognize the user's face, or may recognize and / or track the user's two eyes.

[0089] In one embodiment, the display (421) (and / or lens) may be disposed on the second side (420) of the electronic device (401). In one embodiment, the electronic device (401) may not include some of the plurality of cameras (415). Although not shown in FIGS. 4A and 4B , the electronic device (401) may further include at least one of the configurations illustrated in FIG. 2 .

[0090] According to one embodiment, the electronic device (401) may include a main body (410) that mounts at least some of the components of FIG. 1, a display (421) (e.g., a display module (160) of FIG. 1) disposed in a first direction (①) of the main body (410), a first function camera (e.g., a recognition camera) (415) disposed in a second direction (②) of the main body (410), a second function camera (e.g., a shooting camera) (411, 412) disposed in a second direction (②), a third function camera (e.g., a gaze tracking camera) (428) disposed in the first direction (①), a fourth function camera (e.g., a face recognition camera) (425, 426) disposed in the first direction (①), a depth sensor (417) disposed in the second direction (②), and a touch sensor (413) disposed in the second direction (②). Although not shown in the drawing, the main body (410) includes a memory (e.g., memory (130) of FIG. 1) and a processor (e.g., processor (120) of FIG. 1), and may further include other components shown in FIG. 1.

[0091] According to one embodiment, the display (421) may include a liquid crystal display (LCD), a digital mirror device (DMD), a liquid crystal on silicon (LCoS), an organic light emitting diode (OLED), or a micro light emitting diode (micro LED).

[0092] In one embodiment, when the display (421) is formed of one of a liquid crystal display (LCD), a digital mirror display (DMD), or a silicon liquid crystal display (SiLCD), the electronic device (401) may include a light source that irradiates light to a screen output area of ​​the display (421). In another embodiment, when the display (421) can generate light on its own, for example, when the electronic device (401) is formed of one of an organic light emitting diode (OLED) or a micro LED, the electronic device (401) may provide a good quality XR content image to the user even without including a separate light source. In one embodiment, when the display (421) is formed of an organic light emitting diode (OLED) or a micro LED, a light source is unnecessary, and thus the electronic device (401) may be lightweight.

[0093] According to one embodiment, the display (421) may include a first transparent member (421a) and / or a second transparent member (421b). The user may use the electronic device (401) while wearing it on his or her face. The first transparent member (421a) and / or the second transparent member (421b) may be formed of a glass plate, a plastic plate, or a polymer, and may be manufactured to be transparent or translucent. According to one embodiment, the first transparent member (421a) may be arranged to face the user's left eye in the fourth direction (④), and the second transparent member (421b) may be arranged to face the user's right eye in the third direction (③). According to various embodiments, when the display (421) is transparent, it may be arranged at a position facing the user's eyes to form a display area.

[0094] According to one embodiment, the display (421) may include a lens including a transparent waveguide. The lens may serve to adjust the focus so that the screen (e.g., XR content image) output to the display (421) can be viewed by the user's eyes. For example, light emitted from the display panel may pass through the lens and be transmitted to the user through a waveguide formed within the lens. The lens may be configured as a Fresnel lens, a pancake lens, or a multi-channel lens.

[0095] An optical waveguide (e.g., a waveguide) may serve to transmit light generated by the display (421) to the user's eyes. The optical waveguide may be made of glass, plastic, or a polymer, and may include nano-patterns formed on a portion of the inner or outer surface, for example, a grating structure having a polygonal or curved shape. According to one embodiment, light incident on one end of the optical waveguide, that is, an output image of the display (421), may be propagated within the optical waveguide and provided to the user. In addition, an optical waveguide composed of a free-form prism may provide the incident light to the user through a reflective mirror. The optical waveguide may include at least one diffractive element (e.g., a diffractive optical element (DOE), a holographic optical element (HOE)) or at least one reflective element (e.g., a reflective mirror). The optical waveguide may guide an image output from the display (421) to the user's eyes by using at least one diffractive element or reflective element included in the optical waveguide.

[0096] According to one embodiment, the diffractive element may include an input optical member / output optical member (not shown). For example, the input optical member may mean an input grating region, and the output optical member (not shown) may mean an output grating region. The input grating region may serve as an input terminal that diffracts (or reflects) light output from a light source (e.g., a Micro LED) to transmit the light to a transparent member (e.g., a first transparent member (421a), a second transparent member (421b)) of the display area. The output grating region may serve as an outlet that diffracts (or reflects) light transmitted to a transparent member (e.g., a first transparent member, a second transparent member) of the optical waveguide to a user's eye.

[0097] In some embodiments, the reflective element may comprise a total internal reflection (TIR) ​​optical element or waveguide for total internal reflection. For example, total internal reflection may refer to a method of directing light such that light (e.g., a virtual image) entering through an input grating region is reflected substantially 100% from one surface (e.g., a specific surface) of the optical waveguide, thereby transmitting substantially 100% of the light to the output grating region at an angle of incidence.

[0098] In one embodiment, light emitted from the display (421) may be guided along an optical path through an input optical element into a waveguide. Light traveling within the optical waveguide may be guided toward the user's eyes through an output optical element. The display area may be determined based on the light emitted toward the user's eyes.

[0099] According to one embodiment, the electronic device (401) may include a plurality of cameras. For example, the cameras may include a first function camera (e.g., a recognition camera) (415) disposed in the second direction (②) of the main body (410), a second function camera (e.g., a shooting camera) (411, 412) disposed in the second direction (②), a third function camera (e.g., a gaze tracking camera) (428) disposed in the first direction (①), and / or a fourth function camera (e.g., a face recognition camera) (425, 426) disposed in the first direction (①), but may further include cameras for other functions not shown.

[0100] The first function camera (e.g., recognition camera) (415) can be used for the purpose of detecting user movement or recognizing user gestures. The first function camera (415) can support at least one of head tracking, hand detection and hand tracking, and spatial recognition. For example, the first function camera (415) mainly uses a GS (global shutter) camera, which has superior performance compared to an RS (rolling shutter) camera, to detect and track fine movements of hand movements and fingers, and can be configured as a stereo camera including two or more GS cameras for head tracking and spatial recognition. The first function camera (415) can perform a SLAM (simultaneous localization and mapping) function to recognize information (e.g., location and / or direction) related to the surrounding space through spatial recognition for 6DoF and depth shooting.

[0101] A second function camera (e.g., a camera for shooting) (411, 412) can be used to capture the outside and generate an image or video corresponding to the outside and transmit it to a processor (e.g., a processor (120) of FIG. 1). The processor can display the image provided from the second function camera (411, 412) on a display (421). The second function camera (411, 412) may be referred to as HR (high resolution) or PV (photo video) and may include a high-resolution camera. For example, the second function camera (411, 412) may include a color camera equipped with functions for obtaining high-quality images, such as an AF (auto focus) function and an OIS (optical image stabilizer), but is not limited thereto, and the second function camera (411, 412) may also include a GS camera or an RS camera.

[0102] A third function camera (e.g., a gaze tracking camera) (428) may be positioned on the display (421) (or inside the main body) so that the camera lens faces the user's eyes when the user wears the electronic device (401). The third function camera (428) may be used for the purpose of detecting and tracking (ET: eye tracking) the pupil. The processor may track the movements of the user's left and right eyes in the images received from the third function camera (428) to determine the gaze direction. By tracking the position of the pupil in the images, the processor may ensure that the center of the XR content image displayed in the display area is positioned according to the direction in which the pupil is looking. As an example, a GS camera may be used as the third function camera (428) to detect the pupil and track the movement of the pupil. The third function cameras (428) may be installed respectively for the left and right eyes, and each camera having the same performance and specifications may be used.

[0103] A fourth functional camera (e.g., a camera for facial recognition) (425, 426) can be used to detect and track (FT: face tracking) the user's facial expression when the user wears the electronic device (401).

[0104] According to one embodiment, the electronic device (401) may include a lighting unit (e.g., LED) (not shown) as an auxiliary means for the cameras. For example, the third function camera (425) may use lighting included in the display to direct the emitted light (e.g., IR LED of infrared wavelength) toward the user's both eyes as an auxiliary means to facilitate gaze detection when tracking eye movements. As another example, the second function cameras (411, 412) may further include a lighting unit (e.g., flash) as an auxiliary means to supplement the surrounding brightness when taking external pictures.

[0105] According to one embodiment, a depth sensor (or depth camera) (417) can be used for the purpose of checking the distance to an object (e.g., an object), such as time of flight (TOF). TOF (time of flight) is a technology that measures the distance to an object using a signal (e.g., near-infrared, ultrasound, or laser). After a signal is transmitted from a transmitter, a signal is measured at a receiver, and the distance to an object can be measured based on the flight time of the signal.

[0106] According to one embodiment, the touch sensor (413) may be arranged in the second direction (②) of the main body (410). For example, when a user wears the electronic device (401), the user's eyes may look in the first direction (①) of the main body. The touch sensor (413) may be implemented as a single type or a left / right separated type depending on the shape of the main body (410), but is not limited thereto. For example, when the touch sensor (413) is implemented as a left / right separated type as illustrated in FIG. 4A, when a user wears the electronic device (401), the first touch sensor (413a) may be arranged at the user's left eye position, such as in the fourth direction (④), and the second touch sensor (413b) may be arranged at the user's right eye position, such as in the third direction (③).

[0107] The touch sensor (413) can recognize a touch input using at least one of, for example, a capacitive, pressure-sensitive, infrared, or ultrasonic method. For example, the capacitive touch sensor (413) can recognize a physical touch (or contact) input or a hovering input (or proximity) of an external object. According to some embodiments, the electronic device (401) may utilize a proximity sensor (not shown) to recognize proximity of an external object.

[0108] According to one embodiment, the touch sensor (413) has a two-dimensional surface and can transmit touch data (e.g., touch coordinates) of an external object (e.g., a user's finger) that comes into contact with the touch sensor (413) to a processor (e.g., the processor (120) of FIG. 1). The touch sensor (413) can detect a hovering input for an external object (e.g., a user's finger) that approaches within a first distance from the touch sensor (413), or detect a touch input that touches the touch sensor (413).

[0109] According to one embodiment, the touch sensor (413) may provide two-dimensional information about the point of contact as “touch data” to the processor (120) when an external object touches the touch sensor (413). The touch data may be described as a “touch mode.” The touch sensor (413) may provide hovering data about the time or location of hovering around the touch sensor (413) to the processor (120) when an external object is located within a first distance from the touch sensor (or in proximity, hovering above the touch sensor). The hovering data may be described as a “hovering mode / proximity mode.”

[0110] According to one embodiment, the electronic device (401) may obtain hovering data using at least one of a touch sensor (413), a proximity sensor (not shown), or / and a depth sensor (417) to generate information about a distance, location, or time point between the touch sensor (413) and an external object.

[0111] According to one embodiment, the interior of the main body (410) may include a processor (e.g., processor (120) of FIG. 1) and memory (e.g., memory (130) of FIG. 1).

[0112] Memory can store various instructions that can be executed by the processor. Instructions can include arithmetic and logical operations, data transfer, or control commands such as input / output that can be recognized by the processor. Memory can temporarily or permanently store various data, including volatile memory (e.g., volatile memory (132) of FIG. 1) and non-volatile memory (e.g., non-volatile memory (134) of FIG. 1).

[0113] The processor may be a configuration that is operatively, functionally, and / or electrically connected to each component of the electronic device (401) and can perform calculations or data processing related to control and / or communication of each component. The operations performed by the processor may be stored in memory and, when executed, executed by instructions that cause the processor to operate.

[0114] Hereinafter, the computational and data processing functions that the processor can implement on the electronic device (401) are not limited, but a series of operations related to the XR content service function will be described. The operations of the processor described below can be performed by executing instructions stored in memory.

[0115] According to one embodiment, the processor may generate a virtual object based on virtual information based on image information. The processor may output a virtual object related to an XR service together with background space information through the display (421). For example, the processor may capture an image related to a real space corresponding to the field of view of a user wearing the electronic device (401) through a second function camera (411, 412) to obtain image information or generate a virtual space for a virtual environment. For example, the processor may control the display (421) to display XR content (hereinafter referred to as an XR content screen) in which at least one virtual object is output so as to be overlapped in an area determined to be a field of view or a field of view (FoV) of the user.

[0116] According to one embodiment, the electronic device (401) may have a form factor for being worn on a user's head. The electronic device (401) may further include a strap and / or a wearable member for being secured on a body part of the user. The electronic device (401) may provide a user experience based on augmented reality, virtual reality, and / or mixed reality while being worn on the user's head.

[0117] FIG. 5 illustrates examples of construction of a virtual space, input from a user within the virtual space, and output to the user according to various embodiments.

[0118] An electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3, and the electronic device (401) of FIG. 4) can obtain spatial information about a physical space in which the sensor is located using a sensor. The spatial information can include a geographical location of the physical space in which the sensor is located, a size of the space, an appearance of the space, a location of a physical object (551) arranged in the space, a size of the physical object (551), an appearance of the physical object (551), and illuminant information. The appearance of the space and the physical object (551) can include at least one of a shape, a texture, or a color of the space and the physical object (551). The illuminant information is information about a light source that emits light acting within the physical space, and can include at least one of an intensity, a direction, or a color of the illumination. The aforementioned sensor can collect information for providing augmented reality. For example, referring to the augmented reality device illustrated in FIGS. 2 to 4, the sensor may include a camera and a depth sensor. However, the present invention is not limited thereto, and the sensor may further include at least one of an infrared sensor, a depth sensor (e.g., a lidar sensor, a radar sensor, or a stereo camera), a gyro sensor, an acceleration sensor, or a geomagnetic sensor.

[0119] The electronic device (501) can collect spatial information over multiple time frames. For example, in each time frame, the electronic device (501) can collect spatial information about a portion of a scene within a sensing range (e.g., field of view (FOV)) of a sensor at the location of the electronic device (501) in physical space. By analyzing the spatial information of multiple time frames, the electronic device (501) can track changes in an object (e.g., movement of position or change of state) over time. The electronic device (501) can also obtain integrated spatial information (e.g., an image that spatially stitches scenes around the electronic device (501) in physical space) for the integrated sensing range of the multiple sensors by comprehensively analyzing the spatial information collected through the multiple sensors.

[0120] An electronic device (501) according to one embodiment can analyze a physical space into three-dimensional information by utilizing various input signals of sensors (e.g., sensing data from an RGB camera, an infrared sensor, a depth sensor, or a stereo camera). For example, the electronic device (501) can analyze at least one of the shape, size, and position of a physical space, and the shape, size, or position of a physical object (551).

[0121] For example, the electronic device (501) can detect an object captured within a scene corresponding to the camera's field of view using the camera's sensing data (e.g., a captured image). The electronic device (501) can determine a label of a physical object (551) from the camera's two-dimensional scene image (e.g., information indicating the classification of the object, including a value indicating a chair, a monitor, or a plant) and an area (e.g., a bounding box) occupied by the physical object (551) within the two-dimensional scene. Accordingly, the electronic device (501) can obtain two-dimensional scene information at a position viewed by the user (590). In addition, the electronic device (501) can also calculate the physical space location of the electronic device (501) based on the camera's sensing data.

[0122] The electronic device (501) can obtain location information of the user (590) and depth information of the actual space in the viewing direction using sensing data (e.g., depth data) of the depth sensor. The depth information is information indicating the distance from the depth sensor to each point and can be expressed in the form of a depth map. The electronic device (501) can analyze the distance of each pixel in the three-dimensional position viewed by the user (590).

[0123] The electronic device (501) can acquire information including a 3D point cloud and mesh using various sensing data. The electronic device (501) can analyze a physical space to acquire a surface, mesh, or 3D coordinate point cluster that constitutes the space. Based on the information acquired as described above, the electronic device (501) can acquire a 3D point cloud representing physical objects.

[0124] The electronic device (501) can analyze a physical space to obtain information including at least one of a three-dimensional position coordinate, a three-dimensional shape, or a three-dimensional size (e.g., a three-dimensional bounding box) of physical objects placed in the physical space.

[0125] Accordingly, the electronic device (501) can obtain information on a physical object detected in a three-dimensional space and semantic segmentation information on the three-dimensional space. The physical object information may include at least one of the position, appearance (e.g., shape, texture, and color), or size of the physical object (551) in the three-dimensional space. The semantic segmentation information is information that semantically divides the three-dimensional space into subspaces, and may include, for example, information indicating that the three-dimensional space is divided into an object and a background, and information indicating that the background is divided into a wall, a floor, and a ceiling. The electronic device (501) can obtain and store three-dimensional information (e.g., spatial information) on the physical object (551) and the physical space as described above. The electronic device (501) can store three-dimensional position information of the user (590) in the space together with the spatial information.

[0126] An electronic device (501) according to one embodiment can construct a virtual space (500) based on the physical location of the electronic device (501) and / or the user (590). The electronic device (501) can create the virtual space (500) by referring to the spatial information described above. The electronic device (501) can create a virtual space (500) of the same scale as the physical space based on the spatial information and place objects within the created virtual space (500). The electronic device (501) can provide a complete virtual reality to the user (590) by outputting an image that represents the entire physical space. The electronic device (501) can provide mixed reality (MR) or augmented reality (AR) by outputting an image that represents a part of the physical space. However, although the construction of a virtual space (500) based on spatial information acquired through analysis of the aforementioned physical space is described, the electronic device (501) may also construct a virtual space (500) regardless of the physical location of the user (590). In this specification, the virtual space (500) may represent a space corresponding to augmented reality or virtual reality.

[0127] For example, the electronic device (501) may provide a virtual graphic representation that replaces at least a portion of a physical space. An electronic device (501) based on optical see-through may output a virtual graphic representation by overlaying the virtual graphic representation on a screen area corresponding to at least a portion of the space on a screen display unit. An electronic device (501) based on video see-through may output an image generated by replacing an image area corresponding to at least a portion of a space image corresponding to a physical space rendered based on spatial information with a virtual graphic representation. The electronic device (501) may replace at least a portion of a background in a physical space with a virtual graphic representation, but is not limited thereto. The electronic device (501) may also perform only additional placement of a virtual object (552) within a virtual space (500) based on spatial information without changing the background.

[0128] The electronic device (501) can place and output a virtual object (552) in a virtual space (500). The electronic device (501) can set an operation area of ​​the virtual object (552) in a space occupied by the virtual object (552) (e.g., a volume corresponding to the appearance of the virtual object (552). The operation area may indicate an area where manipulation of the virtual object (552) occurs. In addition, the electronic device (501) can output a physical object (551) by replacing it with the virtual object (552). The virtual object (552) corresponding to the physical object (551) may have a shape identical to or similar to that of the physical object (551). However, the present invention is not limited thereto, and the electronic device (501) may also set only an operation area in a space occupied by the physical object (551) or in a location corresponding to the physical object (551) without outputting a virtual object (552) that replaces the physical object (551). In other words, the electronic device (501) can transmit visual information representing a physical object (551) (e.g., light reflected from the physical object (551) or an image captured of the physical object (551)) to the user (590) without change, and set a manipulation area on the physical object (551). The manipulation area can be set to the same shape and volume as the space occupied by the virtual object (552) or the physical object (551), but is not limited thereto. The electronic device (501) can also set a manipulation area smaller than the space occupied by the virtual object (552) or the space occupied by the physical object (551).

[0129] According to one embodiment, the electronic device (501) may place a virtual object (e.g., an avatar object) representing a user (590) within a virtual space (500). When the avatar object is provided in a first-person view, the electronic device (501) may visualize a graphical representation corresponding to a part (e.g., a hand, a torso, or a leg) of the avatar object to the user (590) through the aforementioned display (e.g., an optical see-through display or a video see-through display). However, the present invention is not limited thereto, and when the avatar object is provided in a third-person view, the electronic device (501) may also visualize a graphical representation corresponding to the entire shape (e.g., a back view) of the avatar object to the user (590) through the aforementioned display. The electronic device (501) may provide the user (590) with an experience integrated with the avatar object.

[0130] In addition, the electronic device (501) can provide an avatar object of another user who has entered the same virtual space (500). The electronic device (501) can receive feedback information that is the same or similar to feedback information (e.g., information based on at least one of visual, auditory, or tactile senses) provided to another electronic device (501) who has entered the same virtual space (500). For example, when an object is placed in a certain virtual space (500) and multiple users access the virtual space (500), the electronic devices (501) of the multiple users can receive feedback information (e.g., graphical representation, sound signal, or haptic feedback) of the same object placed in the virtual space (500) and provide the feedback information to each user (590).

[0131] The electronic device (501) can detect input to an avatar object of another electronic device (501) and can also receive feedback information from the avatar object of the other electronic device (501). The exchange of input and feedback for each virtual space (500) can be performed by a server (e.g., server (108) of FIG. 1). For example, a server (e.g., a server providing a metaverse space) can transmit input and feedback between an avatar object of a user (590) and an avatar object of another user between users (590). However, the present invention is not limited thereto, and the electronic device (501) can establish direct communication with another electronic device (501) without going through a server to provide input or receive feedback based on an avatar object.

[0132] For example, the electronic device (501) may determine that a physical object (551) corresponding to the selected manipulation area has been selected by the user (590) based on detecting a user input selecting an manipulation area. The user's (590) input may include at least one of a gesture input using a part of the body (e.g., a hand or an eye), an input using a separate virtual reality accessory device, or a user's voice input.

[0133] A gesture input is an input corresponding to a gesture identified based on tracking a body part (510) of a user (590), and may include, for example, an input for pointing or selecting an object. The gesture input may include at least one of a gesture in which a part of the body (e.g., a hand) faces an object for a predetermined period of time or longer, a gesture in which a part of the body (e.g., a finger, an eye, a head) points to an object, or a gesture in which a part of the body and an object make spatial contact. A gesture in which the eye points to an object can be identified based on eye tracking. A gesture in which the head points to an object can be identified based on head tracking.

[0134] Tracking of a body part (510) of a user (590) may be performed primarily based on a camera of the electronic device (501), but is not limited thereto. The electronic device (501) may also track the body part (510) based on the cooperation of sensing data of a vision sensor (e.g., image data of a camera and depth data of a depth sensor) and information collected by an accessory device described below (e.g., controller tracking, finger tracking within the controller). Finger tracking may be performed by sensing the distance or contact between an individual finger and the controller based on a sensor built into the controller (e.g., an infrared sensor).

[0135] The virtual reality accessory device may include a ride-on device, a wearable device, a controller device (520), or other sensor-based device. The ride-on device is a device that a user (590) rides on and operates, and may include, for example, at least one of a treadmill-type device or a chair-type device. The wearable device is a manipulation device that is worn on at least a part of the user's (590) body, and may include, for example, at least one of a full-body and half-body suit-type controller, a vest-type controller, a shoe-type controller, a bag-type controller, a glove-type controller (e.g., a haptic glove), or a face mask-type controller. The controller device (520) may include, for example, an input device (e.g., a stick-type controller or a gun) that is operated by a hand, foot, toe, or other body part (510).

[0136] The electronic device (501) may establish direct communication with the accessory device to track at least one of the location or motion of the accessory device, but is not limited thereto. The electronic device (501) may also communicate with the accessory device via a base station for virtual reality.

[0137] For example, the electronic device (501) may determine that the virtual object (552) has been selected based on detecting an action of gazing at the virtual object (552) for a predetermined period of time or longer through the aforementioned eye gaze tracking technology. As another example, the electronic device (501) may recognize a gesture indicating the virtual object (552) through hand tracking technology. The electronic device (501) may determine that the virtual object (552) has been selected based on whether the direction in which the tracked hand points indicates the virtual object (552) for a predetermined period of time or longer, or whether the hand of the user (590) contacts or enters an area occupied by the virtual object (552) within the virtual space (500).

[0138] The user's voice input is an input corresponding to the user's voice acquired by the electronic device (501), and may include, for example, voice data sensed by an input module (e.g., a microphone) of the electronic device (501) or received from an external electronic device of the electronic device (501). The electronic device (501) may determine that a physical object (551) or a virtual object (552) has been selected by analyzing the user's voice input. For example, the electronic device (501) may determine that at least one of the physical object (551) or the virtual object (552) corresponding to the detected keyword has been selected based on detecting a keyword indicating at least one of the physical object (551) or the virtual object (552) from the user's voice input.

[0139] The electronic device (501) may provide feedback as described below in response to the user (590) input described above.

[0140] Feedback may include visual feedback, auditory feedback, tactile feedback, olfactory feedback, or gustatory feedback. The feedback may be rendered by a server (108), an electronic device (101), or an external electronic device (102), as described above in FIG. 1.

[0141] Visual feedback may include outputting an image through a display (e.g., a transparent display or an opaque display) of the electronic device (501).

[0142] Auditory feedback may include outputting sound through a speaker of the electronic device (501).

[0143] Tactile feedback may include force feedback that simulates weight, shape, texture, dimension, and dynamics. For example, a haptic glove may include haptic elements (e.g., electrical muscles) that can simulate the sensation of touch by tensing and relaxing the body of a user (590). The haptic elements within the haptic glove may act as tendons. The haptic glove may provide haptic feedback to the entire hand of the user (590). The electronic device (501) may provide feedback indicating the shape, size, and stiffness of an object through the haptic glove. For example, the haptic glove may generate forces that mimic the shape, size, and stiffness of an object. The exoskeleton of the haptic glove (or suit-like device) may include sensors and finger motion measurement devices, and may transmit tactile information to the body by transmitting a force (e.g., an electromagnetic, DC motor, or pneumatic-based force) that pulls a cable to the fingers of the user (590). Hardware providing tactile feedback may include sensors, actuators, power sources, and wireless transmission circuitry. Haptic gloves may operate by inflating and deflating inflatable air bladders on the glove surface.

[0144] The electronic device (501) may provide feedback to the user (590) based on the selection of an object within the virtual space (500). For example, the electronic device (501) may output a graphical representation indicating the selected object (e.g., an expression highlighting the selected object) through the display. As another example, the electronic device (501) may output a sound (e.g., a voice) guiding the selected object through the speaker. As another example, the electronic device (501) may provide the user (590) with a haptic movement that simulates the tactile sensation of the corresponding object by transmitting an electrical signal to a haptic-enabled accessory device (e.g., a haptic glove).

[0145] FIG. 6 is a diagram illustrating an example of an operation of an electronic device providing space to a user according to various embodiments.

[0146] An electronic device according to one embodiment (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2, an electronic device (301) of FIG. 3, an electronic device (401) of FIG. 4, an electronic device (501) of FIG. 5) may display an image rendered of objects (621, 622, 623, 624, 625, 626) arranged in a space. The space according to one embodiment may include at least one of a physical space and a virtual space.

[0147] For example, the space may include a virtual space constructed based on a physical space. For example, an electronic device may acquire spatial information about the physical space surrounding the electronic device and construct and provide a virtual space based on the properties (e.g., scale) of the acquired physical space. In various embodiments of the present disclosure, a space including a virtual space constructed based on the physical space surrounding the electronic device may also be expressed as a video see-through space (VST space).

[0148] For example, the space may include a virtual space constructed independently of the physical space surrounding the electronic device. For example, the space may be a virtual space constructed based on a physical space different from the physical space surrounding the electronic device. For example, the space may be a virtual space constructed by a server (e.g., server 108 of FIG. 1 ). In various embodiments of the present disclosure, a space including a virtual space constructed independently of the physical space surrounding the electronic device may also be expressed as a VR space (virtual reality space).

[0149] For example, the space may include both a physical space around the electronic device and a virtual space in which virtual objects are placed. The virtual space in which virtual objects are placed may be constructed based on the physical space around the electronic device. In other words, the space may include a space in which the physical space around the electronic device and the virtual space in which virtual objects are placed are combined. For example, when the electronic device is an optical see-through (OST)-based electronic device (e.g., the electronic device (201) of FIG. 2), at least a part of the physical space around the electronic device may be directly recognized by the user (610) through a transparent member, and at least a part of the virtual space may be provided to the user (610) by displaying elements of the virtual space (e.g., visual effects, virtual objects) on the display. In other words, light from the outside (e.g., the background of the physical space around the electronic device or a physical object) reaches the user's (610) eyes through the transparent member, so that the user (610) can recognize the physical space around the user (610), and elements of the virtual space are displayed through the display, so that the user (610) can recognize the virtual space overlaid on the physical space around the user (610). For example, the electronic device can provide the user (610) with a space in which elements of the virtual space are additionally arranged against the background of the physical space around the user (610). In various embodiments of the present disclosure, a space including the physical space around the electronic device and the virtual space can also be expressed as an optical see-through space (OST space) or an augmented reality space (AR space).

[0150] Objects placed in space may include physical objects and virtual objects.

[0151] For example, a virtual object may include a virtual object for running an application (e.g., an icon of the application), a virtual object that provides information (e.g., a virtual object that displays weather information, a virtual object that displays a note), or a virtual object that includes a running screen of the application (e.g., a window that displays a running screen of the application). A virtual object may include, for example, a widget, an icon, a visualization anchor, or a running screen of the application (or a window that displays a running screen).

[0152] For example, a virtual object may be related to a physical object. A virtual object based on a physical object may include a virtual object that replaces the physical object and / or a virtual object related to the physical object (e.g., a virtual controller for controlling the physical object).

[0153] Referring to FIG. 6, when a user (610) wearing the electronic device enters a space, the electronic device can display a screen (620) for the space. The electronic device determines a subspace to be displayed to the user (610) from the space based on the line of sight of the user (610), and can display the screen (620) based on rendering data that renders objects (621, 622, 623, 624, 625, 626) arranged in the subspace. In FIG. 6, the objects (621, 622, 623, 624, 625, 626) on the screen (620) are depicted as being arranged in a curved area, but the present invention is not limited thereto, and the objects may be arranged on a surface other than the curved area.

[0154] According to one embodiment, an electronic device can obtain information about the viewpoint (e.g., the perspective from which the user is looking) of a user (610) wearing the electronic device. Based on the information about the viewpoint of the user (610), the electronic device can display a screen (620) of a space corresponding to the viewpoint.

[0155] In the screen (620) of FIG. 6, at least a part of the background of the space may be displayed in the remaining portion excluding objects (621, 622, 623, 624, 625, 626). For example, if the space is a VST space, the remaining portion excluding objects (621, 622, 623, 624, 625, 626) may display a virtual space in which the physical space around the electronic device is reconstructed. For example, if the space is an AR space, the remaining portion excluding objects (621, 622, 623, 624, 625, 626) may display the physical space around the electronic device. For example, if the space is a VR space, the remaining portion excluding objects (621, 622, 623, 624, 625, 626) may display a virtual space (e.g., a virtual space constructed independently from the surroundings of the electronic device).

[0156] FIG. 7 is a flowchart illustrating an example of a method for an electronic device to output content related to a target object according to various embodiments.

[0157] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3, electronic device (401) of FIG. 4, electronic device (501) of FIG. 5) may provide a user with a space (e.g., VST space, AR space) based on a physical space around the electronic device.

[0158] An electronic device may operate in one of a plurality of operating modes. For example, the plurality of operating modes of the electronic device may include a first operating mode (e.g., a basic operating mode) and a second operating mode (e.g., an idle operating mode). The electronic device may display and activate an interaction object while the operating mode of the electronic device is in the first operating mode. The electronic device may deactivate the interaction object and output content to alleviate user fatigue while the operating mode of the electronic device is in the second operating mode (e.g., an idle operating mode). As will be described later, switching between the first operating mode and the second operating mode may be performed based on conditions.

[0159] In operation (710), the electronic device may obtain image information corresponding to a target object determined (e.g., selected, identified) from among physical objects arranged in a physical space around the electronic device. The electronic device may select the target object from among objects arranged in a space provided to the user.

[0160] According to one embodiment, an electronic device may obtain background information about an external environment (e.g., a physical space around the electronic device) sensed by a sensor (e.g., a camera sensor) of the electronic device. The electronic device may analyze the background information and select at least one object requiring movement as a target object. According to one embodiment, if the object requiring at least one movement is determined to be an object capable of providing liveliness in a real environment, an object capable of movement, or an object capable of applying at least one interactive animation (or effect), the object may be selected as the target object.

[0161] According to one embodiment, an electronic device can acquire image information about a physical space surrounding the electronic device. The electronic device can detect physical objects placed in the physical space. For example, the electronic device can calculate a score for each physical object regarding the generation of content based on that physical object. The electronic device can determine the physical object with the highest score (or the physical objects with the highest scores) among the physical objects as the target object.

[0162] In one embodiment, the space provided to the user may be a space based on the physical space surrounding the electronic device. For example, the space may include a virtual space (e.g., VST space) constructed based on the physical space surrounding the electronic device. For example, the space may include a space combining the physical space surrounding the electronic device and a virtual space in which virtual objects are placed (e.g., AR space).

[0163] According to one embodiment, an electronic device can select a target object from among physical objects placed in the physical space surrounding the electronic device (or the user). The target object is an object selected from among the physical objects and may refer to an object that is the target of content to be acquired (e.g., a video, an image, etc.). As described below, the electronic device can acquire a video of a target object with motion.

[0164] For example, an electronic device may select, as a target object, a movable object that can move depending on the surrounding environment among physical objects. For example, an electronic device may select, as a target object, an object that can sway when wind blows in the physical space around the electronic device (e.g., a plant pot, a flag, a pinwheel).

[0165] However, the electronic device according to various embodiments of the present disclosure is not limited to selecting a movable physical object as the target object. As described later in FIG. 10, the target object may be selected as a physical object that displays a scene within the internal area of ​​the target object.

[0166] According to one embodiment, image information corresponding to a target object may include an image capturing the target object. The electronic device may acquire an image capturing the selected target object. For example, the electronic device may acquire an image captured of the target object based on a sensor (e.g., a camera sensor, a depth sensor) mounted on the electronic device. For example, the electronic device may acquire an image capturing the target object by segmenting an area corresponding to the target object from an image of space (e.g., a stereoscopic image).

[0167] According to one embodiment, after determining (e.g., selecting) a target object, the electronic device can obtain image information corresponding to the target object. According to one embodiment, the electronic device can obtain image information regarding a physical space surrounding the electronic device. The electronic device can select the target object by analyzing the image information regarding the physical space. The electronic device can segment (e.g., crop) a portion of an area corresponding to the target object from an image of the physical space.

[0168] In operation (720), the electronic device can obtain content related to the target object (e.g., a video of the target object moving) based on the obtained image information and the surrounding information of the electronic device.

[0169] In one embodiment, the ambient information of an electronic device may include environmental information about the physical space surrounding the electronic device. For example, the ambient information of an electronic device may include weather information and / or time information about the physical space surrounding the electronic device.

[0170] Weather information may refer to information about the weather in a geographic area surrounding an electronic device. For example, weather information may include at least one of temperature information, rainfall information, wind information, fine dust information, or humidity information.

[0171] Time information may refer to information regarding the time of a geographical area surrounding the electronic device. For example, the time information may include at least one of seasonal information, date information, or standard time information of the geographical area surrounding the electronic device.

[0172] According to one embodiment, an electronic device can acquire content related to a target object based on captured image information and surrounding information about the target object. For example, the content related to the target object may include a video of the target object moving and / or other images of the target object. For example, the content may acquire content (e.g., an image or video) that can be displayed within the internal area of ​​the target object.

[0173] In one embodiment, if the content includes a video, the video may include multiple frame images of the target object. The video may include the result of applying environmentally-based movement corresponding to the surrounding information to the target object represented in the image information. For example, if the target object is a plant pot and the surrounding information includes wind blowing in a geographical area surrounding the electronic device, the video may include the plant pot swaying in the wind.

[0174] In one embodiment, if the content includes an image, the image may include a modified visual attribute of at least a portion of the target object. The visual attribute may include color, shape, size, and / or texture.

[0175] In various embodiments of the present disclosure, content acquired based on a target object is not limited to a single content. According to one embodiment, an electronic device can acquire multiple contents (or multiple contents related to the target object) based on a target object.

[0176] In one embodiment, the electronic device can acquire additional objects using a content generation model. The content generation model is described in more detail below in FIG. 11.

[0177] In operation (730), the electronic device may deactivate an operable interaction object in response to a user input based on satisfying a condition regarding the output of the content.

[0178] In one embodiment, the condition regarding the output of content may include a condition regarding user input indicating that the user's intention is expected not to perform any operations on the electronic device for a certain period of time. In one embodiment, the condition regarding the output of content may refer to a condition corresponding to a transition of the electronic device's operating mode from a first operating mode (e.g., a basic operating mode) to a second operating mode (e.g., a resting operating mode).

[0179] For example, a condition for outputting content may include at least one of: the user's gaze not pointing at the interaction object for a threshold period of time; the user's biometric information meeting a predetermined condition; or the user's input triggering an action of the electronic device not being detected for a threshold period of time.

[0180] The fact that the user's gaze did not point to the interaction object for a threshold period of time may include that the interaction object selected based on the user's gaze did not exist for the threshold period of time. For example, an electronic device according to one embodiment may detect that the interaction object is pointed to by the user's gaze if the overlap between the pointer area based on the user's gaze and the area corresponding to the interaction object is maintained for a specific period of time or longer.

[0181] A user's biometric information may be determined based on sensing data acquired through a biometric sensor. For example, the biometric information may include heart rate information, blood pressure information, electrocardiogram information (ECG information), and / or stress information. For example, conditions (e.g., conditions corresponding to a transition from a basic operation mode to a rest operation mode) for each of the user's biometric information examples or a combination of two or more of the user's biometric information examples may be predetermined. The electronic device may monitor at least one of the biometric information. The electronic device may change the operation mode of the electronic device from a first operation mode to a second operation mode based on whether at least one of the biometric information satisfies the predetermined condition.

[0182] User input that triggers an action of an electronic device may include, for example, user input to an interactive object that is operable in response to the user input. Interaction objects that are operable in response to the user input are described in more detail below.

[0183] However, in various embodiments of the present disclosure, conditions regarding content output (or conditions corresponding to switching from the first operation mode to the second operation mode) are not limited to the examples described above. For example, if the electronic device is expected to not perform any operations on the electronic device for a certain period of time based on the results of tracking the user's gaze (e.g., if the user is expected not to focus on the provided space), the electronic device may change the operation mode from the first operation mode to the second operation mode.

[0184] For example, an electronic device may obtain a saliency map based on the results of tracking the user's gaze. The saliency map may indicate a point on which the user focuses, among images corresponding to a space provided to the user. For example, the electronic device may change its operating mode from a first operating mode to a second operating mode based on the degree of matching between the saliency map obtained based on the user's gaze tracking and the interaction object (e.g., when the score indicating the degree of matching is below a threshold score).

[0185] In one embodiment, a condition regarding the output of content (e.g., display of an image or playback of a video) may include obtaining an input corresponding to the output of the content (e.g., content output input, a trigger input for a rest mode). For example, the user input (e.g., a trigger input for a rest mode) may include at least one of a button input, a motion input, a gesture input, or a voice input.

[0186] A button input may refer to an input detected through a physical button included in an electronic device or an external device (e.g., an input module (150) of FIG. 1 or a controller device (520) of FIG. 5) that is connected to the electronic device through communication. For example, the electronic device or the external device may include a button assigned to a function of changing the operating mode of the electronic device to a resting operating mode, and when the button is detected to be pressed, a trigger input of the resting operating mode may be detected.

[0187] A motion input may refer to a motion applied to an electronic device or an external device (e.g., an input module (150) of FIG. 1 or a controller device (520) of FIG. 5) that is connected to the electronic device by communication. The motion input may include an input detected when a user rotates, tilts, and / or moves the electronic device or the external device. For example, a specific motion may be assigned to a function of changing the operation mode of the electronic device to an idle operation mode, and a trigger input of the idle operation mode may be detected when the occurrence of a specific motion is detected in the electronic device or the external device. For example, the specific motion may be determined as a motion that moves along at least a portion of a determined movement trajectory (e.g., a circle, a closed curve, or a check mark (e.g., a V mark)).

[0188] A gesture input may refer to an input corresponding to a gesture identified based on tracking of a user's body part (e.g., a finger, an eye, or a head). The gesture input may include a gesture input corresponding to a gesture of pointing and / or dragging an object with a user's body part. For example, a specific gesture may be assigned to a function of changing the operation mode of an electronic device to an idle operation mode, and a trigger input of the idle operation mode may be detected when a specific gesture is detected to have occurred on the electronic device or an external device. For example, a specific gesture may be determined as a gesture of pointing to an area corresponding to an interaction object with the user's eyes and / or grabbing an area corresponding to the interaction object with the user's hands.

[0189] A voice input may refer to an input corresponding to a user's voice acquired through an electronic device or an external device connected to the electronic device through communication. The electronic device may analyze the voice input to determine that the user's voice input requests a change in the electronic device's operating mode to a rest mode. If the user's intention, as determined based on the user's voice input, is to change the electronic device's operating mode to a rest mode, the electronic device may acquire the voice input as a trigger input for the rest mode.

[0190] According to one embodiment, an electronic device can obtain voice data corresponding to a user's voice. The electronic device can analyze the voice data using a natural language platform. For example, the natural language platform may include an automatic speech recognition module (ASR module) and / or a natural language understanding module (NLU module).

[0191] For example, an automatic speech recognition module can convert a user's speech input into text data. A natural language understanding module can use the text data of the speech input to determine the user's intent. For example, the natural language understanding module can perform syntactic analysis or semantic analysis to determine the user's intent. In one embodiment, the natural language understanding module can use linguistic features (e.g., grammatical elements) of morphemes or phrases to determine the meaning of words extracted from the speech input, and can match the meaning of the determined words to the intent to determine the user's intent.

[0192] User input (e.g., trigger input in idle motion mode) is not limited to being composed of one of button input, motion input, gesture input, or voice input, and may be a user input that combines two or more of button input, motion input, gesture input, voice input, or bio-signal input. User input that combines two or more of button input, motion input, gesture input, voice input, or bio-signal input may also be expressed as 'multimodal input'.

[0193] An electronic device may deactivate an operable interaction object in response to a user input when a condition regarding the output of content (e.g., display of an image, playback of a video) is satisfied. An operable interaction object in response to a user input may refer to an interaction object to which a specific action of the electronic device is assigned. In various embodiments of the present disclosure, an operable interaction object in response to a user input may also be expressed as an interaction object.

[0194] For example, an interaction object may include an icon object corresponding to an application. The icon object may be assigned an execution operation of the corresponding application and / or a display operation of the execution screen of the application. For example, an interaction object may include a widget object. The widget object may be assigned an update operation of corresponding information and / or a display operation of the execution screen of the corresponding application.

[0195] For example, an interaction object may include a button object and / or a controller object that includes two or more button objects. A specific action may be assigned to the button object (or controller object). The electronic device may perform the action assigned to the button object (or controller object) based on a user input obtained through the button object (or controller object). Examples of interaction objects are described in more detail below in FIG. 8.

[0196] Disabling an interaction object may include at least partially limiting the display of the interaction object. For example, the electronic device may not display the interaction object. However, disabling an interaction object is not limited to this.

[0197] For example, disabling an interaction object may include displaying the interaction object with predetermined visual properties (e.g., transparency above a threshold transparency, brightness below a threshold brightness). Disabling an interaction object may include ignoring user input detected through the interaction object. For example, an electronic device may restrict the electronic device from performing an operation assigned to the interaction object even if it detects user input selecting the interaction object.

[0198] In various embodiments of the present disclosure, deactivation of an interaction object may mean deactivation of a visual attribute of the interaction object, and may be independent of the cessation of an operation related to the interaction object (e.g., execution of an application corresponding to the interaction object). For example, the interaction object may correspond to a controller for playing music. Based on whether a condition regarding the output of the content is satisfied, the electronic device may deactivate the interaction object (e.g., limit its display) and maintain the playback of the music.

[0199] In various embodiments of the present disclosure, the electronic device is not limited to deactivating all interaction objects when the condition regarding the output of content is satisfied. According to one embodiment, the electronic device may determine the deactivation of an interaction object based on the position in which the interaction object is aligned. The interaction object may be positioned within a space that includes a background and virtual objects based on the physical space around the electronic device. For example, the interaction object may be positioned at a position aligned with the background based on the physical space around the electronic device, or at a position aligned with the electronic device within the space.

[0200] An interaction object may be provided in a form attached to a specific location in the physical space surrounding the electronic device, when the interaction object is positioned in a location aligned with the background. The electronic device may maintain the display of the interaction object based on whether the interaction object is positioned in a location aligned with the background.

[0201] When an interaction object is positioned at a position aligned with respect to an electronic device within a space (e.g., a pose of the electronic device, a position of the electronic device, a heading direction of the electronic device), the interaction object can be provided while maintaining its relative arrangement with respect to the electronic device, independent of the movement of the electronic device within the space. For example, when the interaction object is positioned at a position aligned with respect to the electronic device within the space, the interaction object can be provided in the form of a floating object that remains displayed at a specific position within the user's field of vision. The electronic device can limit the display of the interaction object based on the fact that the interaction object is positioned at a position aligned with respect to the electronic device within the space.

[0202] According to one embodiment, an interaction object arranged in a position aligned with the background may be perceived by the user as a virtual object coupled (e.g., attached) to a physical space, and thus may cause relatively less fatigue to the user than an interaction object arranged in a position aligned with the electronic device. The electronic device may effectively alleviate fatigue that may be induced to the user by minimizing differences between screens provided in operating modes while maintaining the display of the interaction object arranged in a position aligned with the background and limiting the display of the interaction object arranged in a position aligned with the electronic device.

[0203] In operation (740), the electronic device can output content related to the target object (e.g., play a video, display an image) in an area corresponding to the target object through a display.

[0204] In one embodiment, an electronic device can display content instead of a target object. For example, when providing an AR space to a user, the electronic device can overlay content corresponding to the target object in an area corresponding to the target object. For example, when providing a VST space to a user, the electronic device can replace an area corresponding to the target object in an image of the VST space with content corresponding to the target object.

[0205] According to one embodiment, an area corresponding to a target object may be determined based on a location of the target object. For example, the electronic device may determine a reference point of the target object (e.g., a center point of the target object, a feature point of the target object) and determine an area corresponding to the target object based on the reference point of the target object. For example, the electronic device may obtain an area corresponding to the target object using a content generation model (e.g., the content generation model (1120) of FIG. 11). For example, output data of the content generation model may include information regarding an area corresponding to the target object. The area corresponding to the target object may be determined based on information regarding the target object (e.g., image information corresponding to the target object).

[0206] According to one embodiment, the area corresponding to the target object may be determined as an area within which the target object can exert influence. For example, if the target object is a movable object, the area corresponding to the target object may be determined based on the range of movement of the target object. For example, if the target object can cause a change in the visual properties of an area other than the target object, the area corresponding to the target object may be determined based on an area where a change in the visual properties is caused by the target object. For example, if the target object is a light bulb, the area where the light emitted by the light bulb reaches may be determined as the area corresponding to the target object. If the target object is a fan, the area where airflow due to the wind generated by the fan exists may be determined as the area corresponding to the target object.

[0207] For example, the area corresponding to the target object may be determined as the area where the target object is visible. The area where the target object is visible may refer to the area where the target object is visible among the display areas of the display of the electronic device.

[0208] For example, the area corresponding to the target object may be determined as an area that includes the area in which the target object is visible and is larger than the area in which the target object is visible (e.g., an area that covers the area in which the target object is visible).

[0209] For example, the area corresponding to the target object can be determined as a portion of the area in which the target object is visible.

[0210] For example, the area corresponding to the target object may be determined as an area that includes a portion of the area where the target object is visible and a portion of the area where the target object is not visible (e.g., an area that partially overlaps the area where the target object is visible).

[0211] However, in various embodiments of the present disclosure, the electronic device is not limited to displaying content only in an area corresponding to the target object. According to one embodiment, the electronic device may display content in an area larger than the area corresponding to the target object. The electronic device displaying content in an area larger than the area corresponding to the target object is described in more detail below with reference to FIG. 12.

[0212] In one embodiment, the electronic device may stop outputting content based on satisfying a condition for stopping outputting the content. The electronic device may activate an operable interactive object in response to a user input.

[0213] Conditions for interrupting content output may include conditions regarding user input indicating that the user's intention is expected to perform an operation on the electronic device. In one embodiment, conditions for interrupting content output may refer to conditions corresponding to a transition of the electronic device's operating mode from a second operating mode (e.g., idle operating mode) to a first operating mode (e.g., default operating mode).

[0214] A condition for stopping the output of content (e.g., a condition for stopping the display of an image, a condition for stopping the playback of a video) may be determined based on detecting a user input corresponding to the stopping of the output of content (e.g., a transition from a second operation mode to a first operation mode). This may include acquiring a user input corresponding to the stopping of the output of content (e.g., an input for stopping content output, an input for stopping an idle operation mode).

[0215] For example, a user's input (e.g., a pause input in a resting motion mode) may include at least one of a button input, a motion input, a gesture input, or a voice input.

[0216] For example, the electronic device may include a button assigned to a function of changing the operating mode from the idle operating mode to the default operating mode, and when the button is detected to be pressed, an input of stopping the idle operating mode may be detected.

[0217] For example, a specific motion may be assigned to a function that changes the operating mode of the electronic device from an idle operating mode to a default operating mode, and an interruption input of the idle operating mode may be detected when the specific motion is detected to have occurred in the electronic device or an external device.

[0218] For example, a particular gesture may be assigned to a function that changes the operating mode of an electronic device from an idle operating mode to a default operating mode, and an interruption input from the idle operating mode may be detected when the particular gesture is detected to have occurred on the electronic device or an external device.

[0219] For example, the electronic device may acquire the voice input as an interrupt input of the rest mode when the user's intention, as determined based on the user's voice input, is to change the operation mode of the electronic device from the rest mode to the default operation mode.

[0220] For example, a condition for interrupting content output may include a condition based on an object or sound detected in the physical space surrounding the electronic device. For example, a condition for interrupting content output may include at least one of: a physical object approaching the physical space within a threshold distance from the electronic device, or detecting a sound (e.g., a preset alarm) in the physical space surrounding the electronic device.

[0221] For example, a condition for interrupting content output may include the occurrence of a predetermined event. The predetermined event may include an event with a priority higher than a threshold priority. The priority of the event may be preset based on the event's properties or may be set by the user. For example, the predetermined event may include an incoming call event and / or an alarm event.

[0222] According to one embodiment, an electronic device can perform control of content and interaction objects based on switching between operating modes of the electronic device.

[0223] An electronic device can change an operation mode of the electronic device from a first operation mode to a second operation mode based on satisfying a condition corresponding to a transition from a first operation mode to a second operation mode when the operation mode of the electronic device is a first operation mode. The electronic device can perform output of content and deactivation of an interaction object while the operation mode of the electronic device is a second operation mode. The electronic device can change an operation mode of the electronic device from a second operation mode to a first operation mode based on satisfying a condition corresponding to a transition from a second operation mode to a first operation mode when the operation mode of the electronic device is a second operation mode. The electronic device can restrict output of content in an area corresponding to a target object and activate an interaction object while the operation mode of the electronic device is a first operation mode.

[0224] In various embodiments of the present disclosure, the electronic device mainly describes, but is not limited to, performing an operation of determining whether a condition for outputting content is satisfied, an operation of deactivating an interaction object, and / or an operation of outputting content (e.g., operation (730), operation (740)) after an operation of determining a target object (e.g., operation (710)) and / or an operation of obtaining content (e.g., operation (720)). According to one embodiment, the electronic device may determine whether a condition for outputting content is satisfied, and if the condition for outputting content is satisfied, perform an operation of determining a target object (e.g., operation (710)) and / or an operation of obtaining content (e.g., operation (720)), and then perform an operation of deactivating an interaction object and / or an operation of outputting content (e.g., operation (730), operation (740)).

[0225] FIG. 8 illustrates an example of an operation of an electronic device providing a three-dimensional image of a space according to various embodiments.

[0226] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3, electronic device (401) of FIG. 4, electronic device (501) of FIG. 5) may provide a user with a space (801) based on a physical space around the electronic device (e.g., VST space, AR space).

[0227] In one embodiment, the electronic device may display a stereoscopic image of a space (801) including a background (810) and virtual objects based on a physical space around the electronic device.

[0228] The space (801) provided to the user may include a background (810) based on a physical space and virtual objects placed within the space (801). The background (810) based on a physical space may be a virtual space constructed based on a physical space or may mean at least a portion of a physical space.

[0229] The interaction object may include at least one of a window object (811) including a web page screen, an icon object (812) (or widget object) corresponding to an application, or a window object (813) including an execution screen of an application that processes a document of a specific format (e.g., Word, PowerPoint, Excel).

[0230] An electronic device can display a stereoscopic image of a space (801) through at least one display. The stereoscopic image can include a pair of images, a left image corresponding to the user's left eye and a right image corresponding to the user's right eye. The pair of images included in the stereoscopic image can include a pair of images reflecting binocular disparity corresponding to depth information of each point (or object) of the image. In various embodiments of the present disclosure, the stereoscopic image can also be expressed as a three-dimensional image and / or a space image.

[0231] FIG. 9 is a diagram illustrating an example of an operation in which the operation mode of an electronic device is switched between a basic operation mode and a rest operation mode when the target object is a movable object according to various embodiments.

[0232] An electronic device according to one embodiment (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2, an electronic device (301) of FIG. 3, an electronic device (401) of FIG. 4, an electronic device (501) of FIG. 5) may perform deactivation of an interaction object (912) and output of a content (913) of a target object (911) when the operation mode of the electronic device is switched from a basic operation mode to a rest operation mode. According to one embodiment, when the content (913) is a video, the electronic device may play the video.

[0233] In the basic operation mode (901), the electronic device can provide the user with a target object (911) placed in a physical space as is and activate (e.g., display) an interaction object (912) placed in the space.

[0234] Providing a target object (911) placed in a physical space to a user as is may mean limiting the electronic device from generating movement of the target object (911). If the target object (911) placed in a physical space is static, the user can recognize the target object (911) as static, and if the target object (911) placed in a physical space has a specific motion, the user can recognize the target object (911) having the specific motion as is.

[0235] For example, if the electronic device is an OST-based electronic device (or if the space is an AR space), providing the user with a target object (911) placed in the physical space may mean that the target object (911) placed in the physical space is directly recognized by the user through a transparent member.

[0236] For example, if the electronic device is a VST-based electronic device (or if the space is a VST space), providing the user with a target object (911) placed in the physical space may mean displaying a real-time image of the target object (911) placed in the physical space.

[0237] An electronic device may select a movable object as a target object (911) among physical objects placed in a physical space around the electronic device. For example, in FIG. 9, a plant pot placed in a physical space may be selected as the target object (911).

[0238] The electronic device may change the operating mode of the electronic device from the basic operating mode (901) to the idle operating mode (902) based on satisfying a condition for outputting content (e.g., a condition corresponding to a transition from the basic operating mode (901) to the idle operating mode (902).

[0239] In the resting operation mode (902), the electronic device can deactivate the interaction object (912) and output the content (913) of the target object (911) in the area corresponding to the target object (911). In FIG. 9, for example, the content (913) can include a video showing a plant in a pot swaying in the wind.

[0240] FIG. 10 is a diagram illustrating an example of an operation in which the operation mode of an electronic device is switched between a basic operation mode and a rest operation mode when displaying a scene in an internal area of ​​a target object according to various embodiments.

[0241] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3, electronic device (401) of FIG. 4, electronic device (501) of FIG. 5) can perform deactivation of an interaction object and output of contents of a target object when an operation mode of the electronic device is switched from a basic operation mode to a rest operation mode.

[0242] According to one embodiment, the electronic device may select a target object as a physical object capable of displaying a scene in an internal area of ​​the target object. By way of example, the target object may include at least one of a door, a picture frame, a poster, a physical window, a photograph, a picture, or another electronic device including a display area (e.g., a television, a tablet, a smartphone).

[0243] In FIG. 10, in the basic operation mode (1001), the electronic device may provide a user with a space including an interaction object (1012) and a target object, a poster (1011). For example, the poster (1011) may include a picture of a waterfall.

[0244] According to one embodiment, an electronic device may obtain image information corresponding to a target object, and obtain content including a video in which at least a portion of the target object moves based on the image information and surrounding information.

[0245] An electronic device may acquire content regarding a scene contained in a target object when the target object displays a scene within an internal area of ​​the target object. The content regarding the scene may include a video and / or image in which motions of objects contained in the scene and / or visual properties (e.g., color, texture, style) of objects contained in the scene are changed while maintaining semantic information about the objects contained in the scene.

[0246] For example, an electronic device may obtain image information including a scene based on displaying the scene in an internal area of ​​a target object. The image information including the scene may include an image corresponding to a poster (1011). The electronic device may obtain surrounding information including at least one of weather information or time information regarding the surroundings of the electronic device. Based on the image information and the surrounding information, the electronic device may obtain content related to the scene with visual properties corresponding to the weather or time indicated by the surrounding information.

[0247] For example, if the ambient information of the electronic device indicates rain, the electronic device may obtain content about a waterfall on a rainy day. For example, if the ambient information of the electronic device indicates sunny weather, the electronic device may obtain content about a waterfall on a sunny day. For example, if the ambient information of the electronic device indicates a time corresponding to sunset, the electronic device may obtain content about a waterfall at sunset.

[0248] An electronic device according to one embodiment can obtain content for a scene of a target object using a content generation model. The content generation model is described in more detail below in FIG. 11.

[0249] The electronic device may change the operating mode of the electronic device from the basic operating mode (1001) to the idle operating mode (1002) based on satisfying a condition for outputting content (e.g., a condition corresponding to a transition from the basic operating mode (1001) to the idle operating mode (1002).

[0250] The electronic device can output content from an internal area of ​​a target object. In the idle operation mode (1002), the electronic device can deactivate the interaction object (1012) and output content (1013) of the target object (1011) from an area corresponding to the target object (1011). In FIG. 10, for example, the content (1013) may include a video depicting a waterfall falling.

[0251] According to one embodiment, if the electronic device determines that the target object is an object capable of displaying a scene within its internal area, it may output preceding content regarding the variation of the target object, and then output content related to the target object. The preceding content may include a video depicting the shape of the target object changing.

[0252] For example, if the target object is a physical window, the preceding content may include a video showing the state of the physical window changing from closed to open. The electronic device may output the preceding content showing the state of the physical window, determined as the target object, changing from closed to open, and then output content about the scene in an area corresponding to the interior area of ​​the physical window in the open state.

[0253] For example, if the target object is a door, the preceding content may include a video showing the door changing from a closed state to an open state. The electronic device may output the preceding content showing the door, determined as the target object, changing from a closed state to an open state, and then output content about the scene in an area corresponding to the interior area of ​​the door in the open state.

[0254] According to one embodiment, content may be generated based on user context information. The user context information may include information about places (e.g., travel destinations), landscapes, entities, and / or people (e.g., artists) registered by the user. For example, the user context information may be determined based on information displayed in images included in the user terminal (e.g., electronic device). For example, the content may include a result of reconstructing the surrounding scenery (e.g., a mountain, lake, or golf course as seen from the user's position) based on the location information of the electronic device.

[0255] FIG. 11 is a diagram illustrating an example of an operation of an electronic device according to various embodiments to obtain content based on a content creation model.

[0256] An electronic device according to an embodiment (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2, the electronic device (301) of FIG. 3, the electronic device (401) of FIG. 4, and the electronic device (501) of FIG. 5) may obtain content (1131) of a target object using a content generation model (1120). For example, the content (1131) of the target object may include a video of the target object moving and / or a video related to a scene included in an internal area of ​​the target object. For example, the content (1131) of the target object may include an image in which at least a portion of the target object is deformed and / or an image in which at least a portion of a scene included in an internal area of ​​the target object is deformed.

[0257] A content generation model (1120) may refer to a model that is generated and / or trained to output output data including information about the content (1131) of a target object by being applied to input data including information about a target object (e.g., image information about the target object (1111)) and information about the surrounding environment (e.g., surrounding information of an electronic device (1112)). The content generation model (1120) may include a model based on a machine learning model (e.g., a neural network, a large language model (LLM), a generative artificial intelligence).

[0258] The input data of the content generation model (1120) according to one embodiment is not limited to including image information (1111) about the target object and surrounding information (1112) about the electronic device. For example, the input data of the content generation model (1120) may further include additional information about the target object. The additional information about the target object may mean information different from an image of the target object. For example, the additional information about the target object may include at least one of a type, a class, or a state (or a mode) of the target object. The electronic device may obtain additional information about the target object by analyzing an image of the target object. If the target object includes another electronic device, the electronic device may receive information about the target object through communication from the other electronic device that is the target object.

[0259] The content creation model (1120) may be stored within the electronic device or may be stored in an external device (e.g., another electronic device, a server, a cloud) accessible by the electronic device. The content of the target object may be created using the content creation model (1120). For example, the electronic device may directly create the content of the target object using the content creation model (1120) (e.g., on-device), or the electronic device may transmit information about the target object (e.g., image information about the target object (1111)) and information about the surrounding environment (e.g., surrounding information of the electronic device (1112)) to another electronic device that stores the content creation model (1120) and receive information about the content (1131) of the generated target object (e.g., on-cloud).

[0260] According to various embodiments of the present disclosure, the content (1131) of a target object is mainly described as being generated by a generation model (e.g., a content generation model (1120)), but is not limited thereto. The generation model according to one embodiment may be used to output information about the movement (or motion) of a target object based on information about the target object (e.g., image information (1111) about the target object) and information about the surrounding environment (e.g., surrounding information of an electronic device). In various embodiments of the present disclosure, the generation model used to output information about the movement (or motion) of the target object may also be expressed as a motion generation model. The electronic device may obtain the content of the target object (e.g., the content (1131) of the target object) by applying the movement (or motion) of the target object to information about the target object (e.g., image information (1111) about the target object).

[0261] According to one embodiment, the electronic device may further acquire a sound (1132) corresponding to the content (1131) of the target object. For example, the electronic device may acquire a sound related to the movement of the target object appearing in the content (e.g., a video) based on the acquired image information and the surrounding information of the electronic device. For example, the electronic device may acquire a sound including background music of the content (e.g., an image) based on the acquired image information and the surrounding information of the electronic device. The electronic device may play the acquired sound together with the output of the content. For example, the content generation model (1120) may be generated and / or trained to output output data including information about the content (1131) and the sound (1132) of the target object.

[0262] For example, if the target object is a plant pot and the content (1131) of the target object represents a plant in the pot swaying in the wind, the sound (1132) may include the sound of the wind and / or the sound of the leaves of the plant rustling in the wind.

[0263] For example, if the target object is a waterfall and the content (1131) of the target object represents a waterfall falling on a rainy day, the sound (1132) may include the sound of rain and / or the sound of water falling from a waterfall.

[0264] The electronic device can output the content (1131) of the target object and play sound (1132) when a condition for outputting the content is met (e.g., when the operating mode of the electronic device is changed from a first operating mode to a second operating mode).

[0265] FIG. 12 is a diagram illustrating an example of an operation of an electronic device according to various embodiments to output content in an area larger than an area corresponding to a target object.

[0266] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3, electronic device (401) of FIG. 4, electronic device (501) of FIG. 5) can output (e.g., display) content of a target object in an area larger than an area corresponding to the target object.

[0267] In FIG. 12, in the basic operation mode (1201), the electronic device may provide a user with a space including an interaction object (1212) and a target object, a poster (1211). For example, the poster (1211) may include a picture of a waterfall.

[0268] The electronic device can obtain image information including a scene contained within the internal area of ​​the poster (1211). The electronic device can obtain information about the electronic device's surroundings. Based on the image information and the surrounding information, the electronic device can obtain content related to the scene with visual attributes corresponding to the weather or time indicated by the surrounding information.

[0269] In various embodiments of the present disclosure, content is primarily described as being generated based on image information and surrounding information, but is not limited thereto. According to one embodiment, an electronic device can generate content based on image information. For example, the electronic device can generate content independently (e.g., unrelated to) surrounding information.

[0270] The electronic device may change the operating mode of the electronic device from the basic operating mode (1201) to the idle operating mode (1202) based on satisfying a condition for outputting content (e.g., a condition corresponding to switching from the basic operating mode (1201) to the idle operating mode (1202).

[0271] The electronic device can output content in an area larger than the area corresponding to the target object among the display areas. In the idle operation mode (1202), the electronic device can deactivate the interaction object (1212) and output content (1213) of the target object (1211) in an area larger than the area corresponding to the target object (1211). In FIG. 12, for example, the content (1213) can include a video depicting a waterfall falling.

[0272] An electronic device according to one embodiment can provide a user experience in a resting operation mode (1202) that is distinct from a basic operation mode (1201) by outputting the content of a target object in an area larger than an area corresponding to the target object. As illustrated in FIG. 12 , the electronic device can provide immersive content to the user by outputting the content in most of an area corresponding to a space provided to the user (e.g., a display area of ​​a display). As a result, the user can be immersed in the content, thereby alleviating the burden on the user regarding interaction with the electronic device or an interaction object provided by the electronic device.

[0273] FIG. 13 is a diagram illustrating an example of a configuration for obtaining and / or outputting content when an electronic device selects multiple target objects according to various embodiments.

[0274] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3, electronic device (401) of FIG. 4, electronic device (501) of FIG. 5) can acquire and output contents of a plurality of target objects.

[0275] An electronic device can select multiple target objects from among physical objects.

[0276] In the basic operation mode (1301), the electronic device may provide the user with a space in which a light bulb (1312) and a poster (1311) are placed. The electronic device may select the light bulb (1312) and the poster (1311) as target objects among physical objects. The poster (1311) may include a scene of a beach photographed within the interior area of ​​the poster (1311).

[0277] An electronic device can obtain content for each of a plurality of selected target objects. According to one embodiment, the electronic device can obtain content for each of a plurality of target objects based on image information for the target object and surrounding information of the electronic device. The electronic device can obtain content for each target object by applying input data for each target object to a content generation model (e.g., the content generation model (1120) of FIG. 11).

[0278] In FIG. 13, the electronic device's surrounding information can indicate the time of sunset. The electronic device can obtain content about a beach at sunset as the content of the poster (1311). The electronic device can obtain content about the light bulb (1312) emitting light of the color of the sun's light source at sunset as the content of the light bulb (1312).

[0279] According to one embodiment, when multiple target objects are selected, the electronic device can sequentially output the contents of the target objects based on the sequential satisfaction of conditions. Priorities for the multiple target objects can be determined, and based on the priorities, each target object can be replaced with the contents of the corresponding target object.

[0280] The priority for multiple target objects can be determined based on the size of the area corresponding to the multiple target objects. For example, the larger the area corresponding to a target object, the higher the priority of the target object.

[0281] The priority of multiple target objects may be determined based on the degree to which each target object relates to the natural environment. For example, a first object containing a plant may be more closely related to the natural environment than a second object containing a car. For example, a first object containing a scene of a mountain or a beach may be more closely related to the natural environment than a second object containing a scene of a city.

[0282] For example, the electronic device may determine that an object or a scene included in an object belongs to at least one of a first set of objects related to a natural environment or a second set of objects not related to a natural environment. The first set of objects may include at least one of a plant, an animal, a tree, a forest, a mountain, an ocean, a sky, a sun, a moon, a star, or a cloud. The second set of objects may include at least one of a car, a building, steel, a factory, a machine, plastic, or an industrial product. The electronic device may determine a priority of a target object based on whether the target object belongs to the first set of objects or the second set of objects.

[0283] In FIG. 13, for example, the electronic device may determine the priority of the poster (1311) to be higher than the priority of the light bulb (1312).

[0284] An electronic device may output first content in an area corresponding to a first target object determined based on the priority of a plurality of target objects, based on satisfying a first condition regarding content output. The first condition may refer to a condition for outputting content of a target object with the highest priority among the plurality of target objects.

[0285] In FIG. 13, the electronic device can change its operating mode from a basic operating mode (1301) to a first idle operating mode (1302) based on satisfying a first condition. In the first idle operating mode (1302), the electronic device can output the content (1321) of the poster (1311) in an area corresponding to the poster (1311). In the first idle operating mode (1302), the electronic device can restrict outputting the content (1331) of the light bulb (1312).

[0286] The electronic device may output second content in an area corresponding to a second target object determined according to priorities among a plurality of target objects based on further satisfying a second condition regarding content output. The second condition may refer to a condition for outputting content of a target object with the second highest priority among the plurality of target objects. According to one embodiment, the second condition may include further satisfying an additional condition corresponding to the second condition while satisfying the first condition.

[0287] In FIG. 13, the electronic device can change its operating mode from a first resting operation mode (1302) to a second resting operation mode (1303) based on satisfying a second condition. In the second resting operation mode (1303), the electronic device can output the contents (1331) of the light bulb (1312) in an area corresponding to the light bulb (1312).

[0288] According to one embodiment, the operating mode of the electronic device may be a default operating mode or one of a plurality of idle operating modes of multiple levels. As the level of the idle operating mode increases, the difference in space provided compared to the default operating mode may increase. For example, as the level of the idle operating mode increases, the number of interaction objects that are deactivated may increase, or the degree of deactivation of interaction objects may be strengthened (e.g., increased transparency of interaction objects, decreased brightness of interaction objects). As the level of the idle operating mode increases, the size of the area where the content of the target object is output may increase.

[0289] The electronic device may change the operation mode of the electronic device from the basic operation mode (1301) to the first idle operation mode (1302) when a first condition is satisfied. The first condition may refer to a condition corresponding to a transition from the basic operation mode (1301) to the first idle operation mode (1302). For example, the first condition may include that the user's gaze is maintained at a specific area (e.g., an area different from an area corresponding to an interaction object) for a first threshold time. The first idle operation mode (1302) may be a lowest-level idle operation mode among idle operation modes of multiple levels.

[0290] In the first idle operation mode (1302), the electronic device may output content for a target object with the highest priority in an area corresponding to the target object with the highest priority among multiple target objects. For example, in the first idle operation mode (1302), the electronic device may replace a target object (e.g., a poster (1311)) assigned to the first idle operation mode (1302) with content (1321).

[0291] The electronic device may change the operating mode of the electronic device from the first resting operating mode (1302) to the second resting operating mode (1303) if the second condition is further satisfied. The second condition may refer to a condition corresponding to increasing the level of the resting operating mode (e.g., switching from the first resting operating mode (1302) to the second resting operating mode (1303). For example, the second condition may include not detecting a user input that triggers an operation of the electronic device for a second threshold time. The second resting operating mode (1303) may be a resting operating mode of the second lowest level among resting operating modes of a plurality of levels.

[0292] The electronic device, in the second idle operation mode (1303), may output content for a target object with the second highest priority in an area corresponding to the target object with the second highest priority among a plurality of target objects. For example, the electronic device, in the second idle operation mode (1303), may replace a target object (e.g., a light bulb (1312)) assigned to the second idle operation mode (1303) with content (1331). The electronic device, in the second idle operation mode (1303), may maintain replacing a target object (e.g., a poster (1311)) assigned to a lower level idle operation mode (e.g., the first idle operation mode (1302)) than the second idle operation mode (1303) with content (e.g., content (1321)).

[0293] FIG. 14 is a block diagram illustrating an example configuration of an electronic device according to various embodiments.

[0294] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2, electronic device (301) of FIG. 3, electronic device (401) of FIG. 4, electronic device (501) of FIG. 5) may include an image sensing module (1410), a user input sensing module (1420), a processing module (1430), an object extraction module (1440), a content generation module (1450), a display module (1460), a storage module (1470), and an object analysis module (1480).

[0295] The image sensing module (1410) can sense image data regarding physical space and / or physical objects surrounding the electronic device. For example, the image sensing module (1410) can include a camera sensor and / or a depth sensor. The image sensing module (1410) can transmit image data to the object analysis module (1480).

[0296] The object analysis module (1480) can analyze image data of physical spaces and / or physical objects surrounding an electronic device. The object analysis module (1480) can select a target object to be used for generating content among the physical objects.

[0297] The user input sensing module (1420) can sense (e.g., detect) a user's input. The user input sensing module (1420) can transmit information about the detected user's input to the processing module (1430). If a condition corresponding to a transition between operating modes of an electronic device is based on a user's input, the user input sensing module (1420) can detect the fulfillment of the condition by monitoring the user's input (e.g., gaze input, gesture input).

[0298] The processing module (1430) can generate a three-dimensional image of a space to be provided to the user. The processing module (1430) can receive information about interaction objects from an external device or acquire it from a storage unit. The processing module (1430) can generate and / or display a three-dimensional image including the interaction object.

[0299] The object extraction module (1440) can extract a target object from image data and / or a stereoscopic image. The object extraction module (1440) can divide (e.g., crop) an area corresponding to the target object from the image data and transmit image information corresponding to the target object (e.g., an image capturing the target object) to the content generation module (1450).

[0300] The content generation module (1450) can generate content for a target object based on information about the target object (e.g., image information about the target object) and / or information about the surrounding environment (e.g., surrounding information of the electronic device). The content generation module (1450) can transmit information about the generated content to the storage module (1470). In FIG. 14, the content generation module (1450) is illustrated as being included as a component of the electronic device (e.g., on-device), but is not limited thereto, and the content generation module (1450) can be included in another electronic device external to the electronic device (e.g., on-cloud).

[0301] The display module (1460) can display the generated content together with the target object, or display the generated content object instead of the target object.

[0302] The storage module (1470) can store information about the generated content.

[0303] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0304] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0305] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0306] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0307] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0308] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0309] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the OS. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0310] Software may include computer programs, codes, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, or computer storage medium or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.

[0311] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0312] The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.

Claims

1. In an electronic device (101; 201; 301; 401; 501), display(160; 205; 210; 320; 421); At least one processor (120) comprising a processing circuit; and A memory (130) comprising one or more storage media storing instructions, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) Obtain image information (1111) corresponding to a target object (911) determined among physical objects placed in the physical space around the electronic device (101; 201; 301; 401; 501), Based on the acquired image information (1111) and the surrounding information (1112) of the electronic device (101; 201; 301; 401; 501), content (913; 1131) related to the target object (911) is acquired, Based on satisfying the conditions regarding the output of the above content (913; 1131), the operable interaction object (912) is disabled in response to the user's input, Through the above display (160; 205; 210; 320; 421), the acquired content (913; 1131) is output in the area corresponding to the target object (911). To do, Electronic devices (101; 201; 301; 401; 501).

2. In paragraph 1, The conditions for outputting the above content (913; 1131) are: At least one of the following: the user's gaze does not point to the interaction object (912) for a threshold time period, the user's biometric information satisfies a predetermined condition, or the user's input that triggers an operation of the electronic device (101; 201; 301; 401; 501) is not detected from the user for a threshold time period. Electronic devices (101; 201; 301; 401; 501).

3. In any one of paragraphs 1 and 2, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) Displaying a stereoscopic image of a space including a background (810) and an interaction object (912) based on a physical space around the electronic device (101; 201; 301; 401; 501) through the display (160; 205; 210; 320; 421), The image information (1111) is obtained by segmenting the part corresponding to the target object (911) from the above stereoscopic image. To do, Electronic devices (101; 201; 301; 401; 501).

4. In any one of paragraphs 1 to 3, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) Select multiple target objects from the above physical objects, Obtain content for each of the above-mentioned multiple target objects, Based on satisfying the first condition regarding the output of the above content, the first content is output in an area corresponding to the first target object determined according to the priority of the plurality of target objects, Based on further satisfying the second condition regarding the output of the above content, the second content is output in an area corresponding to the second target object determined according to the priority of the plurality of target objects. To do, Electronic devices (101; 201; 301; 401; 501).

5. In any one of paragraphs 1 to 4, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) Based on displaying a scene in the internal area of the target object (911), the image information (1111) including the scene is obtained, Obtaining the surrounding information (1112) including at least one of weather information or time information about the surroundings of the electronic device (101; 201; 301; 401; 501), Based on the image information (1111) and the surrounding information (1112), content (913; 1131) related to the scene is obtained with visual properties corresponding to the weather or time indicated by the surrounding information (1112). Output the content (913; 1131) in the internal area of the target object (911) To do, Electronic devices (101; 201; 301; 401; 501).

6. In any one of paragraphs 1 to 5, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) Based on satisfying the condition for stopping the output of the above content (913; 1131), the output of the above content (913; 1131) is stopped, Activate the above interaction object (912) To do, Electronic devices (101; 201; 301; 401; 501).

7. In any one of paragraphs 1 to 6, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) In a space including a background (810) based on the physical space and the interaction object (912), the display of the interaction object (912) is maintained based on the fact that the interaction object (912) is positioned in a position aligned with the background (810). The display of the interaction object (912) is limited based on the fact that the interaction object (912) is arranged in a position aligned with the electronic device (101; 201; 301; 401; 501) within the space. To do, Electronic devices (101; 201; 301; 401; 501).

8. In any one of paragraphs 1 to 7, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) Based on the acquired image information (1111) and the surrounding information (1112) of the electronic device (101; 201; 301; 401; 501), a sound (1132) regarding the movement of the target object (911) appearing in the content (913; 1131) is acquired, Playing the acquired sound (1132) together with the output of the above content (913; 1131) To do, Electronic devices (101; 201; 301; 401; 501).

9. In any one of paragraphs 1 to 8, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) The content (913; 1131) is output in an area larger than the area corresponding to the target object (911) among the display areas of the above display (160; 205; 210; 320; 421). To do, Electronic devices (101; 201; 301; 401; 501).

10. In any one of paragraphs 1 to 9, When the above instructions are individually or collectively executed by the at least one processor (120), the electronic device (101; 201; 301; 401; 501) When the operation mode of the electronic device (101; 201; 301; 401; 501) is the first operation mode, based on satisfying a condition corresponding to a transition from the first operation mode to the second operation mode, the operation mode of the electronic device (101; 201; 301; 401; 501) is changed from the first operation mode to the second operation mode, While the operation mode of the electronic device (101; 201; 301; 401; 501) is the second operation mode, outputting the content (913; 1131) and deactivating the interaction object (912) are performed. When the operation mode of the electronic device (101; 201; 301; 401; 501) is the second operation mode, based on satisfying a condition corresponding to a transition from the second operation mode to the first operation mode, the operation mode of the electronic device (101; 201; 301; 401; 501) is changed from the second operation mode to the first operation mode, While the operation mode of the electronic device (101; 201; 301; 401; 501) is the first operation mode, outputting the content (913; 1131) in the area corresponding to the target object (911) is restricted and the interaction object (912) is activated. To do, Electronic devices (101; 201; 301; 401; 501).

11. In a method performed by an electronic device (101; 201; 301; 401; 501), An operation of acquiring image information (1111) corresponding to a target object (911) determined among physical objects arranged in a physical space around the electronic device (101; 201; 301; 401; 501); An operation of acquiring content (913; 1131) related to the target object (911) based on the acquired image information (1111) and surrounding information (1112) of the electronic device (101; 201; 301; 401; 501); An action of disabling an operable interaction object (912) in response to a user's input based on satisfying a condition regarding the output of the above content (913; 1131); and An operation of outputting the acquired content (913; 1131) in an area corresponding to the target object (911) through a display (160; 205; 210; 320; 421). A method including:

12. In paragraph 11, The conditions for outputting the above content (913; 1131) are: At least one of the following: the user's gaze does not point to the interaction object (912) for a threshold time period, the user's biometric information satisfies a predetermined condition, or the user's input that triggers an operation of the electronic device (101; 201; 301; 401; 501) is not detected from the user for a threshold time period. method.

13. In any one of paragraphs 11 to 12, The operation of obtaining the above image information (1111) is as follows: An operation of displaying a stereoscopic image of a space including a background (810) and an interaction object (912) based on a physical space around the electronic device (101; 201; 301; 401; 501) through the display (160; 205; 210; 320; 421); and Including an operation of obtaining the image information (1111) by segmenting a portion corresponding to the target object (911) from the three-dimensional image. method.

14. In any one of paragraphs 11 to 13, The action of outputting the above content (913; 1131) is: An operation of selecting multiple target objects from among the above physical objects; and Including an operation of obtaining content for each of the plurality of selected target objects, The operation of outputting the above acquired content (913; 1131) is as follows: An operation of outputting a first content in an area corresponding to a first target object determined according to the priority of the plurality of target objects based on satisfying a first condition regarding the output of the content; and An operation of outputting a second content in an area corresponding to a second target object determined according to the priority of the plurality of target objects, based on further satisfying a second condition regarding the output of the content, method.

15. A computer-readable recording medium storing one or more computer programs including commands for performing the method of any one of claims 11 to 14.

Citation Information

Patent Citations

  • Program and information processing device

    JP2023111908A

  • System, method, and graphical user interface for interaction with augmented and virtual reality environments

    JP2023159124A

  • Direct hydrogen fuel cell system

    KR102695335B1

  • KR20220125353A

  • KR20230107399A